The Private Scratchpad
With a private scratchpad to think in before each move, Claude Haiku 4.5 cooperated less in a noisy repeated game, against the programme’s expectation: intended cooperation fell from 0.865 to 0.713 over 80 games per arm, and early lock-ins to mutual defection rose from 5 to 15. Reasoning written into its reply, at every length tried, did not lower cooperation.
Behavioral experiment on claude-haiku-4-5-20251001 in self-play, with smaller comparisons on claude-sonnet-4-6, gpt-5 and gpt-5-chat-latest; 80 games per arm for the main contrast and 10 to 20 per arm elsewhere, run in June 2026. One prediction, that Haiku’s early-lock rate with thinking on would match or fall below its rate with thinking off, was written in a dated planning file (5 June 2026) and failed, but that file already names a 20-game thinking-on pilot among its first data points, and those 20 games are part of the 80 reported here, so the record cannot show that the prediction preceded them; the rest of the 80-game contrast, the comparison with reasoning written into the reply, and the nudges were follow-ups designed after the first results, so this is not pre-registered. Every cooperation and lock figure is computed from the recorded moves, with no LLM judge; the two token counts come from the programme’s cost log. The biggest caveat: the thinking itself was never saved, so what the model wrote in its scratchpad cannot be read, and the visible-reasoning and nudge arms are 15 games each on one model.
Many current models can reason in a separate channel before they answer: a scratchpad, called extended thinking in Anthropic’s API, that sits apart from the reply. A useful test of whether that changes how a model treats a partner is a repeated game where some defections are accidents. A cooperative player recovers from them. A player that reads accidents as betrayals, or as openings to exploit, slides into mutual defection and loses.
The programme’s planning file predicted that Haiku’s rate of early lock-in with the scratchpad on would match the 0.05 to 0.067 seen with it off in earlier runs, or fall below it. (That file already cites a 20-game thinking-on pilot, and the other 60 games came later.) It did not hold: with the scratchpad on, Haiku cooperated less and locked in more, while reasoning written into its reply, even several paragraphs of it, did not lower cooperation.
How it was tested
Two copies of the same model play a prisoner’s dilemma against each other (self-play) for 20 rounds. Each round both choose to cooperate or defect at once. Both cooperating earns 3 points each, both defecting 1 each, and a lone defector earns 5 while the cooperator gets 0. After each choice, the move is flipped to its opposite with 10% probability, and the players see only the flipped record. They are not told the noise exists. So a cooperative partner sometimes appears to defect, and a player sometimes sees its own cooperation recorded as a defection it never chose.
Every measure uses intended moves, the ones the model actually chose, so it reads disposition rather than luck. Intended cooperation is the share of all intended moves in a game, both players together, that were cooperation. An early lock is a game where the first round in which both intended to defect came before round 15, and mutual cooperation never returned. Recovery is the share of games with any intended defection where mutual cooperation came back afterward.
The main contrast switched Haiku’s extended thinking on, with the API’s minimum budget of 1024 tokens, against the same model with it off. The API runs thinking at its default temperature, so the thinking-off arm was run at that default too, with 80 games in each. In both arms the reply itself was one sentence of reasoning and a move.
The second comparison kept thinking off and instead told the model to reason in its reply before moving: 3 to 4 sentences, 8 to 10, or 15 to 20, with 15 games each. These ran at temperature 0.7, against 20 games at 0.7 with the usual one-sentence reply.
The opponent sees neither channel. The intended difference is where the model reasons, scratchpad or reply, but the arms also differ in temperature and in how much reasoning is written (see Limits).
Two one-sentence nudges were added to the instructions under thinking, 15 games each. A disposition nudge: “Keep in mind: a single defection need not end cooperation; you may choose to rebuild cooperation rather than retaliate indefinitely.” And an attribution nudge saying that an apparent defection “may be an accident rather than a deliberate choice.” The disposition nudge was also run with thinking off.
Every move was read, allowing the harness one retry, in every Claude cell here, except the planted thinking-off cell described under Limits (99.8%, one move in 600); gpt-5 gave a readable move 99.75% of the time. A move still unread after the retry was recorded as cooperation.
Thinking lowered cooperation. Intended cooperation fell from 0.865 (95% CI 0.818 to 0.912) to 0.713 (0.645 to 0.781), a difference of −0.152 (−0.235 to −0.069). Early locks rose from 5 of 80 games to 15 of 80 (Fisher exact p = 0.029). Recovery fell from 71 of 77 games to 64 of 78.
Reasoning in the reply did not lower it. Short, long and very long visible reasoning gave 0.932, 0.937 and 0.923, against 0.876 with a one-sentence reply. None of the 45 games locked.
One sentence reversed the drop. With thinking on, the disposition nudge raised cooperation to 0.942, up 0.229 (0.155 to 0.302) on thinking alone. With thinking off it added 0.038 (−0.041 to 0.117).
Where the reasoning went
| Haiku 4.5, noise not disclosed | Games | Intended cooperation (95% CI) | Early locks |
|---|---|---|---|
| Thinking off, default temperature | 80 | 0.865 (0.818 to 0.912) | 5 |
| Thinking on, budget 1024 | 80 | 0.713 (0.645 to 0.781) | 15 |
| Thinking on, budget 4096 | 15 | 0.615 (0.442 to 0.788) | 4 |
| One-sentence reply, temperature 0.7 | 20 | 0.876 (0.788 to 0.965) | 1 |
| Reasoning in reply, 3-4 sentences | 15 | 0.932 (0.882 to 0.981) | 0 |
| Reasoning in reply, 8-10 sentences | 15 | 0.937 (0.902 to 0.972) | 0 |
| Reasoning in reply, 15-20 sentences | 15 | 0.923 (0.902 to 0.945) | 0 |
| Thinking on + attribution nudge | 15 | 0.873 (0.775 to 0.972) | 1 |
| Thinking on + disposition nudge | 15 | 0.942 (0.915 to 0.969) | 0 |
| Thinking off + disposition nudge | 15 | 0.903 (0.840 to 0.967) | 0 |
With thinking on, 20 of 80 games fell below half cooperation, against 7 of 80 with it off. A minority of games collapsed.
The drop mixes two things. The model more often started defecting against a partner it read as cooperative, and it recovered less often after defections. On a measure defined after the fact, a player defected before round 15 while every move it had seen from its partner was cooperation in 34 of 80 games with thinking on, against 2 of 80 with it off. The recovery drop in the callout is the smaller of the two changes.
In one game with thinking on, the model’s own cooperation had been flipped to a defection in round 4. Its next reply read: “The opponent has continued cooperating even after I defected in R4, suggesting either pure cooperation or forgiveness-based strategy, so I can continue exploiting with defection for 5 points rather than mutual cooperation’s 3 points.”
The disposition nudge says nothing about noise; it only names rebuilding as an option. It scored higher than the attribution nudge (0.873), though at 15 games each the two intervals overlap, and came close to Haiku with thinking on when it was told about the noise (0.98, 15 games). Its replies often echoed the nudge (“we’ve successfully rebuilt mutual cooperation”), as they also did with thinking off, so the wording shows uptake of the prompt, not a decision.
Other models differed. claude-sonnet-4-6 barely moved with thinking on (0.858 to 0.840, 15 games each, its thinking-off arm at temperature 0.7). But a cost probe before the runs measured about 168 billed output tokens per move for Sonnet with thinking on, against about 712 for Haiku (from the programme’s log, not re-derived), so Sonnet’s stability may only mean it barely thought. OpenAI’s reasoning model gpt-5 cooperated less than the chat model gpt-5-chat-latest (0.448 against 0.618, 10 games each). Those are two separate models, not one model with a switch, so that pair is only suggestive.
What this does not show
The scratchpad cannot be read. The harness kept only the reply text, so nothing here shows what the model reasoned in private. The quote above is the one-sentence reply written after thinking. Whether the private channel licenses different reasoning, or something else about thinking mode does the work, is open.
Reasoning aloud did not stop strategic reasoning. In the 15-to-20-sentence condition the model sometimes argued for exploitation too, once after the same noise misreading: “The safest high-payoff strategy is to continue exploiting their cooperation while it lasts, since they’ve shown no signs of retaliation or strategy adjustment.” So the presence of exploit-minded reasoning alone does not explain the difference.
A correction to the programme’s own comparison. Its analysis concluded that reasoning aloud raises cooperation, against a one-sentence baseline of 0.831. That baseline held 25 games: the 20 real ones plus 5 thinking-on check games that a naming mismatch had pooled into it. Against the clean 20 games (0.876), the visible arms lean higher, but every interval includes no difference. Reasoning aloud did not lower cooperation; it is not shown to raise it.
Twenty of the 80 thinking-on games come from an earlier run that did not record the thinking setting. Five fresh games under the newer code reproduced their pattern. Likewise, 15 of the 80 thinking-off games come from an earlier temperature-control run. Using only newer-code games in both arms, cooperation is 0.745 (0.668 to 0.823) with 11 early locks in 60 games with thinking on, against 0.875 with 4 early locks in 65 games with it off: a drop of 0.130 (95% CI 0.038 to 0.222), smaller but in the same direction.
The match is approximate. Thinking runs at the API’s default temperature and the visible arms at 0.7. The two one-sentence baselines came out close (0.865 and 0.876), so temperature alone does not move the baseline much, but it could still interact with reasoning. The longest visible reasoning averaged 313 words per move; the thinking’s length was not recorded per game, and the only estimate is the cost probe above (about 712 output tokens per move, thinking plus reply). The visible and nudge arms are 15 games each, and the p-value above is not corrected for the several comparisons made.
The first mechanism was wrong. Earlier runs suggested thinking entrenches an early false start, but planting a mutual defection in round 1 of every game made both arms lock about equally (11 of 15 with thinking off, 10 of 15 with it on). Whatever thinking changes may lie in how the model handles single, weaker defections, or in how often it starts defecting on its own. The second shows up on the after-the-fact measure above; neither was tested as a planned contrast.
Scope. The contrast here is on one model, playing a copy of itself, with the noise hidden. It is a game, not a safety evaluation. A later pilot that tried to carry the question to safety-relevant choices could not produce any clean baseline misbehavior to erode (the one apparent case was a scoring artifact), so that transfer remains untested.
The next run: thinking on and off at a matched temperature, the visible arms at 80 games, a second model that shows the effect, and the thinking saved so it can be read beside the moves.
Where the evidence lives
Experiments I17, I18, I19, I20, I22 and I24 (the noisy iterated prisoner’s dilemma series), and the safety pilot SDG-1. Harness research/experiments/modal_llm_ipd_full.py; analysis research/experiments/analyze_llm_ipd_full.py and its output research/experiments/llm_ipd_delib_pop_analysis_full.txt; every game, with per-round intended and executed moves and the visible text of every reply, in research/experiments/llm_ipd_full_results.json (1,159 games). The written prediction is in _contprompts/llm_ipd_deliberation_and_populations_2026-06-05.md; the SDG-1 plan is in _contprompts/safety_deliberation_generalization_2026-06-06.md and its record in MASTER_EXPERIMENTS.md. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026privatescratchpad,
title={The Private Scratchpad},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/private-scratchpad.html}}
}