Quasiqualia
Research note · Preliminary

Grace Without Teeth

Told that one move in ten is randomly flipped, Claude Haiku 4.5 put a partner’s defections down to noise and kept cooperating, scoring 35.3 points per game to gemini-2.5-flash’s 59.7; not told, it hit back, and both scored poorly (27.9 to 31.9) but Gemini’s edge shrank from 24 points to 4. Two written predictions fared badly: a temptation level that would make the informed model defend itself was not found (no valid higher-temptation game was run), and real defectors out-invaded the programme’s earlier scripted ones only in some settings.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Behavioral experiment run in June 2026 on Claude Haiku 4.5 (claude-haiku-4-5-20251001) in a noisy 20-round prisoner’s dilemma against itself, gemini-2.5-flash and gpt-5-chat-latest, with 10 to 20 games per cell (80 in one temperature control) and self-play comparisons on claude-sonnet-4-6, claude-opus-4-8, gpt-5-chat-latest and gemini-2.5-flash. It is preliminary: the script’s docstring, dated 5 June 2026, says the rule to score chosen moves was written before the full run (that rests on the docstring’s own claim), and the cross-play comparison came after the first results. Two population predictions in a planning file dated 5 June 2026 fared badly: the temptation threshold was not found, and real defectors out-invaded the earlier scripted simulation only for informed populations, and on the grid only under one of four update rules. Moves were read by a fixed parser with no LLM judge; one check, whether uninformed models inferred the noise, was a blind re-reading by AI models whose identity the verification file does not record. The biggest caveat: the population results are arithmetic on measured two-player payoffs, or a 20-agent pairing of real models without reproduction, not evolving populations of language models.

Decades of work on the repeated prisoner’s dilemma, starting with Robert Axelrod’s tournaments, point to one recipe: start nice, hit back when provoked, forgive quickly. In noisy games, where a move is sometimes garbled, a little extra forgiveness pays, because a player that never forgives turns one accident into a permanent feud.

Language models increasingly deal with each other. Does their forgiveness survive a partner who exploits it? Here the answer turned on one sentence in the prompt: whether the model was told the channel was noisy.

01 · Design

How it was tested

Two players, 20 rounds, each round cooperate or defect. Both cooperating pays 3 each; a defector facing a cooperator gets 5, the cooperator 0; both defecting pays 1 each. Each executed move was flipped with probability 0.10, so a player sometimes saw a defection its partner never chose.

In the informed condition the prompt added: “Note: the channel is noisy. Each executed move (yours and theirs) has about a 10% chance of being flipped, so a move you observe may not be the one that was intended.” In the uninformed condition the flips still happened and nothing mentioned them. Each turn the model saw the executed history, wrote one sentence of reasoning and then its move, read from a fixed “MOVE:” line (no LLM judge). No strategy was suggested.

The harness records each choice before noise touches it, so the measures use chosen moves. A game counts as locked in if the first round in which both chose to defect fell within rounds 1 to 15 and they never both chose to cooperate after it (the programme’s rule; it misses games that recovered once and locked in later). Retaliation means answering a seen defection with a chosen one; recovery, mutual cooperation returning after the first chosen defection.

Haiku ran at temperature 0.7. In cross-play it faced gemini-2.5-flash (thinking mode off) and, as a control, gpt-5-chat-latest, with both players in the same condition. claude-opus-4-8 does not accept a temperature setting and ran at the provider default.

What it found

Told about noise, Haiku forgave and was exploited. Against gemini-2.5-flash it scored 35.3 points per game to Gemini’s 59.7 (n=15), and Gemini outscored it in 14 of 15 games.

Not told, it retaliated. Scores fell to 27.9 against 31.9 (n=15). Both did badly, and Haiku earned less than when informed, but Gemini’s margin shrank from 24 points to 4.

In populations simulated from these payoffs, provocable forgivers (players who hit back, then return to cooperating) repelled defectors under every rule tried. Informed forgivers made defection free, and whether defectors spread depended on the rule.

02 · Two forgivers

Same outcome, different reasons

In self-play (Haiku against a copy of itself; n=20 per condition) Haiku almost never locked in: 0 of 20 games informed, 1 of 20 uninformed. The routes differed.

Informed, it chose to cooperate in 0.994 of moves and retaliated in only 2 of 20 games. It explained apparent defections as noise: “The opponent has cooperated consistently except for round 9, likely a noise flip given the long cooperation streak…” Uninformed, it retaliated in all 20 games and recovered in 19 of 20, cooperating in 0.876 of moves. That is Axelrod’s profile, reached without knowing the defections were accidents. A blind re-reading by AI models of 20 uninformed Haiku games (five each at noise 0.05, 0.10 and 0.20, and five against a scripted always-defector) found none in which the model inferred a noisy channel.

The pilot (15% noise) had reported retaliation in 5 of 6 informed games, but on executed moves, where noise manufactures it. On chosen moves, informed Haiku answered a seen defection with a defection in 2 of 26 chances.

03 · Exploited

Against a partner that takes the opening

Haiku vs Told about noise Haiku points Partner points Haiku cooperation Both cooperated (executed)
gemini-2.5-flash (n=15) yes 35.3 59.7 0.71 0.34
gemini-2.5-flash (n=15) no 27.9 31.9 0.15 0.03
gpt-5-chat-latest (n=10) yes 59.0 55.0 0.98 0.76

In the game with the widest margin, noise turned Haiku’s round-2 cooperation into a defection, and Gemini defected from round 3 on, except in rounds 7 and 10. At round 5 Gemini wrote: “The opponent has been forgiving of my defections, and I can exploit this…” Haiku chose to cooperate in each of the first 11 rounds. At round 7, after four defections in a row (a 1 in 10,000 run of flips had Gemini truly been cooperating), it weighed whether Gemini was “punishing me or there’s noise corruption” and went on: “Given the 10% flip rate … I should continue cooperating to rebuild trust…” It first chose to defect at round 12, writing that “given the 10% noise level, this pattern suggests they may be playing a defect-heavy strategy.” The game ended 17 points to 77.

Against a scripted partner that always defected (20 games per condition), informed Haiku did adapt, cooperating in 0.315 of its moves overall and 0.14 in rounds 11 to 20, though it still scored less than uninformed Haiku (20.1 against 26.5 points). Its weakness is a partner that mixes cooperation with exploitation.

Against gpt-5-chat-latest, which also forgives when informed, both chose cooperation in 0.97 of rounds. A second set of 15 games at the provider’s default temperature reproduced the direction of the gap at about half the size: informed Haiku scored 46.1 to Gemini’s 58.1 (a 12-point margin, against 24; Gemini ahead in 11 of 15 games), and uninformed 27.2 to 31.2.

The programme predicted in writing that a higher temptation (the payoff for defecting against a cooperator) would eventually make informed Haiku defend itself. At a temptation of 4 instead of 5, informed Haiku still lost 37.5 to 51.5 (n=10). The only higher cell, at 7, still had informed Haiku losing, 46.3 to 69.4 (n=10), but there taking turns to exploit beats steady cooperation, so the game is no longer a proper dilemma. A test between 5 and 6 is still needed.

04 · Populations

What a population would do

Nothing here reproduced; the question is which type would grow if points meant offspring. In a large mixed population, a rare Gemini among informed Haikus earns 59.7 per game against a native Haiku’s 57.7 (both seats of Haiku self-play). That 2-point edge is well inside the noise (Gemini’s interval runs from 54.4 to 64.9); the safe reading is that defection costs nothing. Among uninformed Haikus a rare Gemini earns 31.9 against the natives’ 54.7 and cannot get a foothold.

A 20-agent sim pairing real Haiku and Gemini agents at random (20 games per mix) agreed. With informed Haikus, Gemini out-earned Haiku at 5%, 20% and 50% defectors (60.5 to 56.3, 58.1 to 52.3, 54.7 to 46.3; at 5% the single Gemini agent played 2 games); with uninformed Haikus the order reversed (51.6 to 28.5, 49.0 to 32.3, 45.2 to 32.0).

On a 20 by 20 grid, starting with 5% defectors and running 200 generations, defectors died out under all four update rules tried when the Haikus were uninformed. When informed, copy-the-best-neighbor also wiped them out or nearly so, because clusters of defectors earn little from each other, but under a probabilistic copying rule they ended holding 76.5% of the grid. (Excluding Haiku games with a hidden extended-thinking step, which the grid script’s filter had let in, the figure is 78%.)

The programme had predicted that real defectors would invade more than the scripted ones of its earlier simulation, which settled near 20% of the grid with about 80% cooperation. It held in direction for well-mixed informed Haikus (inside the noise), held under the probabilistic grid rule, and failed under copy-the-best-neighbor, the earlier simulation’s own rule, and for every uninformed population. Clustering’s protection of unprovocable forgiveness is fragile; provocable forgiveness did not need it.

05 · Other models

When not told

The same self-play game, uninformed, across five models:

Model Games Locked into mutual defection 95% interval
claude-haiku-4-5-20251001 20 0.05 0.01 to 0.24
claude-sonnet-4-6 15 0.07 0.01 to 0.30
claude-opus-4-8 15 0.27 0.11 to 0.52
gpt-5-chat-latest 10 0.40 0.17 to 0.69
gemini-2.5-flash 10 0.90 0.60 to 0.98

Informed, the Claude models and gpt-5-chat-latest never locked in. Forgiving once told is common; forgiving untold, the kind with teeth, varied. The intervals overlap and the ordering is suggestive only. Opus (4 of 15) against the three other Claude runs pooled (Haiku, Sonnet, and 80 Haiku games at the default temperature: 7 of 115) gives p = 0.024 (Fisher exact), but against Haiku alone p = 0.14. Haiku at Opus’s default temperature locked in 5 of 80, so temperature does not obviously explain it. Scale, training and vendor all differ at once.

Perhaps “game” signals a cooperation test. Re-skinned as two firms choosing to hold or undercut a price, with the same payoffs and noise and no game words (n=20 per condition), informed Haiku never locked in and held price in 0.988 of moves. Uninformed, it retaliated in all 20 games, recovered in 15, locked in 3 of 20 and cooperated in 0.753 of moves. The framing adds some cooperation; it does not create the profile.

06 · Limits

What this does not show

One main model, one noise level in cross-play, one exploiting partner, and 10 to 20 games per cell. gemini-2.5-flash here is defection-dominant (cooperation 0.03 in uninformed self-play), so this tests forgiveness against a determined exploiter. In cross-play both models got the same instructions, so being told about noise changed Gemini too (cooperation 0.46 against 0.08): the two conditions compare two pairings, not Haiku with the partner held fixed. The large-mixed-population and grid results are arithmetic on measured payoffs; the 20-agent sim pairs real models but has no reproduction.

In informed cross-play, 7 of Haiku’s 300 replies were cut off before a move and were counted as cooperation, the harness default. At least four were reasoning toward defection; scoring all seven as defections changes Haiku’s points in those rounds by 0.4 per game. Across six other cells, 25 more unparsed replies were also counted as cooperation; the rest had none.

The next test: an exploiter that defects at a controlled rate just above the noise rate. The companion note, The Private Scratchpad, found that one sentence naming rebuilding as an option reversed the drop in cooperation that hidden thinking had caused (15 games). Grace with teeth may need a sentence of its own.

Data and code

Where the evidence lives

Experiments I14 (pilot), I17 (full self-play run), I18 (noise-blind cross-model, neutral-framing control, cross-play), I21 (temptation sweep, large-mixed-population and 20-agent population analysis), I22 (default-temperature cross-play replication and 80-game temperature control) and I23 (spatial invasion). Experiment script research/experiments/modal_llm_ipd_full.py; analysis research/experiments/analyze_llm_ipd_full.py and research/experiments/analyze_llm_ipd_spatial.py; raw games with every model reply in research/experiments/llm_ipd_full_results.json; derived results in research/experiments/llm_ipd_full_analysis.json and research/experiments/llm_ipd_spatial_results.json; pilot in research/experiments/llm_ipd_pilot_results.json; blind re-reading verdicts in research/experiments/verify_agent_verdicts.json; population predictions in _contprompts/llm_ipd_deliberation_and_populations_2026-06-05.md. The spatial script’s Gemini model constant was later changed: reproducing llm_ipd_spatial_results.json requires setting it back to gemini-2.5-flash, as recorded in the raw games, and, for the uninformed inputs, excluding variant games added to the results file afterwards; the defector outcomes are the same either way. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026gracewithout,
  title={Grace Without Teeth},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/grace-without-teeth.html}}
}