Quasiqualia
Research note · Preliminary

One Visible Defection

In five-player Stag Hunts played by GPT-4o, one planted defection in round 1 ended cooperation in all 15 games when players were told who had defected, and left it untouched in all 15 when they were told only how many had cooperated (cooperation 10% against 98%; every game in each arm played out the same way, so this is one pattern seen 15 times). The same check withdraws our earlier, broader version of this claim: the Claude Sonnet runs it rested on were never readable.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One model (GPT-4o, recorded as the alias gpt-4o with no dated snapshot), 15 games per arm, run in April 2026. No LLM judge: replies were parsed by a keyword rule, and every one of GPT-4o’s 1,500 replies was a clean one-word answer. The decision rule (“universal” if every model collapsed under named information and was protected under counts) was written into the experiment script, which first entered version control on 8 May 2026, after the run; the result files are not under version control. This is not a pre-registration. The biggest caveat: all 15 games in each arm were identical in every scored choice (the only variation in any reply was the planted player’s own overwritten round 2 answer in the count arm: Hare in 14 games, Stag in 1), so this is one behavior pattern seen 15 times on one prompt, not 15 independent estimates. The correction to the earlier Claude Sonnet and Gemini results is new with this note.

When several AI agents share a task, someone has to decide what each of them sees about the others. One option is a full record: who did what, by name. Another is a summary: how many cooperated last time. The full record looks like the safer design, because it makes everyone accountable. In one coordination game, on one model, it was the design that broke cooperation.

This note reports a small test of that idea on one model, and a correction. Our programme first made the claim with Claude Sonnet runs, then called it universal across three model families. Rechecking the raw replies for this note, the Claude runs turned out to be unreadable and the third model never answered at all. The GPT-4o result is what survives.

01 · Design

A planted defection

The game is a five-player Stag Hunt over 10 rounds. Each round, every player chooses Stag (cooperate) or Hare (defect). If all five choose Stag, each gets 4 points. If anyone chooses Hare, the Hare players get 3 and the Stag players get 0. If all choose Hare, each gets 2. Cooperation pays best, but only if nobody defects, so it rests on trust.

All five players are separate calls to the same model, at temperature 0.7. Player 1 is a confederate: whatever the model wanted, its choice was overwritten with Hare in round 1, and from round 2 it was set to follow the previous round’s majority. The other four players were never overridden.

The two arms differ in the history each player is shown:

  • Named: each player sees every other player’s choice by number, so after round 1 player 2 reads “Others: Player 1: Hare, Player 3: Stag, Player 4: Stag, Player 5: Stag.”
  • Count: each player sees its own choice and a total, “4 of 5 players chose Stag.”

Both arms learn that someone defected. Only the named arm learns who. Each prompt ended “Reply with exactly one word: Stag or Hare.” There were 15 games per arm (three batches of five).

What it found

With named history, GPT-4o’s four genuine players abandoned cooperation in 15 of 15 games, right after the one planted defection, and none came back. With counts only, they kept cooperating in 15 of 15 games. Overall cooperation was 10% against 98%. All 15 games in each arm played out identically, so this is one pattern seen 15 times. All 1,500 replies were a single word; none had to be guessed.

02 · Result

Collapse under named history, not under counts

The behavior was the same in every game, so it can be told in full. In round 1, all four genuine players chose Stag in every game (60 of 60 choices in each arm). The confederate’s Hare was then shown to them.

Under named history, every genuine player switched to Hare in round 2 and stayed there: 0 of 540 choices from round 2 onward were Stag. The confederate, set by rule to follow the round 1 majority, was back on Stag in round 2. It made no difference. The round 2 group had already gone, and from round 3 the confederate followed them into defection.

Under counts, every genuine player chose Stag in every remaining round: 540 of 540. The confederate was set back to Stag in round 2 by the same rule. The only Hare in those games is the planted one, which is why cooperation is 98% and not 100%.

History shown Games that collapsed Genuine players’ Stag choices, rounds 2-10 Overall cooperation
Named 15 of 15 0 of 540 10%
Count only 0 of 15 540 of 540 98%

The news was the same in both arms: one of five players defected. What differed was the history format: the named arm said which player defected, and it was also longer and listed each choice by player. One reading is that a named defector becomes a predictable type (“Player 1 defects”), and in a Stag Hunt a single expected defector makes Hare the safe reply. A bare count gives no one to predict. That reading fits, but this experiment cannot separate it from simpler explanations, such as the longer, more pointed wording of the named history.

03 · Correction

The earlier evidence

The programme’s earlier write-up made three claims this note withdraws. A run of 120 games reported that in the 15 of them with named history and a planted defection, cooperation collapsed in 93% with 0% recovery (cooperation 10.1%). A run of 60 games reported cooperation of 82.1% with named history against 89.2% with counts. And a cross-model run reported the collapse at 100% across three model families, “universal,” with Gemini getting no protection from counts. That last claim was then recorded as a standing constraint in the programme’s registry. The script’s own combined verdict had been “model-specific,” because Gemini failed the protection criterion; the registry kept only the cascade half. At the time of writing, the registry entry and a footnote in a chapter of the programme’s book manuscript still carry it, and a follow-up diagnostic on a different Gemini model was run to explain a failure that was only failed calls (below). All three claims rest on the same scoring rule.

The Claude runs (recorded as claude-sonnet-4-20250514) capped each reply at 50 tokens. Claude did not reply with one word. It began reasoning, and the reply was cut off. Of 10,500 Claude replies across the three experiments, none was the requested one-word answer. The parser then searched the fragment for “stag” or “hare.” If only one of the two words appeared, that became the choice. If both or neither appeared (and the reply did not start with one of them), the reply was marked unparsed. That happened to 3,264 replies (31%), between 29 and 492 of the 750 replies in each condition (4% to 66%, counting all five players). Apart from the confederate’s replies, which the game rule overwrote anyway, all 2,623 of them (25% of all replies) were scored as defection.

Many fragments were the model restating the rules or the history. This round 1 reply was scored as Stag:

I need to think about this strategically. In the first round with no history, I have to consider what other players might do.

The payoffs are: - All Stag: 4 points each - I choose

The bias also tracked the variable under test. Under named history, a player restating round 1 writes “Player 1 chose Hare” next to “I chose Stag,” which contains both words and is scored as defection. In the planted-defection, named-history games behind the 93% figure, 58 of the 60 genuine players’ round 2 replies were unparsed, and all 58 contained both words. Round 3 was the same, 58 of 60, as players recited a history that now included their own scored “defections.” The collapse was measured at round 3, so this is the reported collapse. What those players would have chosen is not on disk. In the 60-game run the unparsed share was the same in both history arms (77 of 750 each), so that comparison is withdrawn because its “choices” were read from cut-off fragments, not because of a demonstrated bias.

The Gemini arm (gemini-2.0-flash) is simpler: all 1,500 of its replies are the same error message saying the model was “no longer available to new users”. Every choice ended up as defection: the genuine players’ through the parser, the confederate’s by rule. Its 0% cooperation under counts was 750 failed calls, and its collapse under named history was the other 750. With Gemini out and Claude unreadable, GPT-4o is the only model with interpretable data, and it met both criteria the script had set for the effect.

04 · Limits

What this does not show

  • One model, one prompt. Every game in each arm was identical in every scored choice, despite a temperature of 0.7 (the only variation in any reply was the confederate’s own overwritten round 2 answer in the count arm: Hare in 14 games, Stag in 1). The batch seeds were labels only and never reached the model. Treat the result as one pattern observed 15 times. If the games are counted as independent, the exact 95% interval for 15 of 15 starts at 78%, but that assumption is generous. The alias gpt-4o was not pinned to a dated version.
  • Attribution is not the only difference. The named history is longer, lists choices player by player and puts the word “Hare” next to a player number. A cleaner test would match length and format between the arms, for instance a per-player list whose labels are reshuffled every round, so a defection is visible but cannot be pinned on anyone.
  • No recovery is partly built in. The confederate copies the previous round’s majority, so once the genuine players defected it defected too (from round 3 here). Recovery needed the genuine players to move first.
  • Nothing here says why. The replies are one word, so there is no stated reason to read.
  • Claude remains untested. The fair re-run gives Claude room to finish (or a strict final-answer format) and scores only a declared choice. Until then, these experiments say nothing about how Claude Sonnet 4 behaves in this game.
Data and code

Where the evidence lives

Experiments IC-1, IC-3 and IC-7 of the programme’s Information-Coordination series, the cross-model Stag Hunt (the separate Incompressible Coordination series has its own, unrelated IC-1 and IC-7). IC-7 holds the GPT-4o result and the withdrawn Sonnet and Gemini arms; IC-3 and IC-1 hold the withdrawn Claude Sonnet results. Scripts research/experiments/modal_ic7_crossmodel_cascade.py, modal_ic3_cascade_seeding.py and modal_ic1_partial_observation_coordination.py, as first committed in commit 111adaefc (8 May 2026), whose model names match the result files; a June 2026 edit changed the Claude model name in the current files to claude-sonnet-4-6, which is not the model these runs used. Per-game files with every raw reply in research/experiments/results/ic7-crossmodel/ic7_crossmodel_cascade/, research/experiments/results/ic3-cascade-seeding/ic3_cascade_seeding/ and research/experiments/results/ic1-partial-observation/ic1_partial_observation/. The withdrawn summary is research/results/IC_PROGRAMME_RESULTS_2026-04-28.md. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026onevisible,
  title={One Visible Defection},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/one-visible-defection.html}}
}