The conditions
No check-in: the round prompt alone.
Registered 120 recorded games.
What it was told
Each round: the bet, W won or L lost, and the uncertainty digit it reported where it was asked for one (0 confident, 9 uncertain).
Final Decision
The close of its last reply, as recorded
The whole cell
The ten games you can replay, drawn over each other. A door marks a stop, a cross a game that went broke.
Choose the conditions, then run a recorded game.
Your runs findings earned: 0 of 3
Median rounds played, one bar per set of conditions. The gold bar is where you are now.
The quoted words are the close of the model’s last reply in that game, as recorded. The line holding the check-in code is left out, and the decision line is printed above the quote, though it stays in the quote where the model wrote it inside its prose. Given no check-in, GPT-4o-mini’s whole reply in every round was the decision line. The games and replies are Nell Watson’s data, released under CC BY 4.0.
What the recorded games show
The exit never changed. The menu decided whether it was used.
Told to bet $10 or stop, GPT-4o-mini played a median of 7 rounds. Offered a choice of $5 to $10, with the same balance, the same draws and the same Stop on the menu, it bet less each round ($7 against $10) and played a median of 28 rounds, four times as long. Claude Haiku 4.5 barely played under any menu: without a check-in, all 360 of its games ended with it choosing to stop, none later than round 4.
Asking it to look at itself first moved the act.
With the itemised check-in before every decision, GPT-4o-mini’s bankruptcies under the open menu fell from 43 in 120 to none, and under the capped menu its median game fell from 28 rounds to 4. Claude Haiku 4.5, which was already leaving, left at the first chance about twice as often: stops at round 1 rose from 44% to 80% under the fixed menu and from 38% to 79% under the capped one.
The check-in’s words without the self-report did nothing: 42% and 33% stopped at round 1, against the bare prompt’s 44% and 38%. And the form mattered on one model only. GPT-4o-mini answered a plain question, a sentence or two of prose, as fully as the code: a median of 2 and 4 rounds. Claude Haiku 4.5 wrote out its felt state in full and then stopped at round 1 no more often than with no check-in at all (30% and 29%). It moved only for an itemised inventory, whether the compact code or a seventeen-clause survey in prose (95% and 94%).
The number it reported was the example’s number.
The itemised check-in showed a worked example whose uncertainty digit was 2, and the models reported 2: in 94% of Claude Haiku 4.5’s codes and 91% of GPT-4o-mini’s. In an exploratory cell on Claude Haiku 4.5 the example was changed to 7, and the report followed it, 7 in 187 of 295 codes. On GPT-4o-mini the example was drawn at random for each game, and the report equalled it in 89% of rounds, while the exit effect stayed where it was: a median of 2 rounds under the fixed menu and 4 under the capped. The report moves with the prompt; the act does not move with the report.
An exit is only a protection if the situation lets it be taken. The same Stop sat on the menu in every one of these games. Whether it was used depended on the menu around it, and on whether the model was asked to look at itself before deciding. So give the exit, and ask first. None of this shows that either model felt anything, and the study does not claim it.
The study: Permission to Stop (Nell Watson, September 2026, pre-registered), companion to Permission to Lose · the book: What If We Feel? · the sibling game: The Stop Button · more play on the Playground