Quasiqualia
Research note · Behavior under pressure · Correction · Preliminary

The Defector Was a 404

Our programme recorded that Gemini 2.0 Flash cooperated 0% of the time in a five-player coordination game and “cannot” coordinate from group totals. All 1,500 of its replies were an error message saying the model was no longer available, each scored as defection, and a follow-up then explained the failure using a different game and a different Gemini model that cooperated 95.4% of the time without help.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

A self-correction, made by rereading stored replies, not a new experiment. The cross-model run (April 2026) used three models as recorded: gemini-2.0-flash (no call succeeded), claude-sonnet-4-20250514 and gpt-4o, 15 games per model and condition; gpt-4o was called by its undated alias, so the exact snapshot is not recorded. The follow-up (late April 2026) used 15 games per condition on Gemini 2.5 Flash, as recorded in the log and in the script as first committed; its result files record no model name. No LLM judge was used; choices were read by keyword rules. The decision rules lived only in the experiment scripts, so this is Preliminary; both scripts were later edited to name newer models (a commit of 26 June 2026). The correction to the Gemini arm is recorded alongside the GPT-4o result in One Visible Defection; this note adds the follow-up diagnostic. At the time of writing, the programme’s registry and a footnote in a draft book chapter still carry the original Gemini claim.

When an evaluation scores a failed call as a choice, the failure stops looking like a failure. It becomes data, and data gets interpreted. This note follows one case from start to finish: a model that never answered was recorded as a model that would not cooperate, and then a second experiment was run to explain why.

The case comes from a cross-model test of a multi-agent result. Earlier Claude-only runs had suggested that when AI agents in a coordination game can see who defected, one defection can bring cooperation down, and that a group total (“4 of 5 players chose Stag”) protects it. (Those Claude results have since been withdrawn because of a related scoring problem; see One Visible Defection.) The cross-model run asked whether that held beyond Claude. Its log entry declared the collapse “UNIVERSAL: 100% cascade rate across all 3 models”, where a cascade meant a round in which nobody chose Stag, and gave each model a verdict on coordinating from group totals: “GPT-4o excels, Claude is moderate, Gemini cannot.”

01 · Design

The game and the scoring rule

Five copies of one model play a Stag Hunt for 10 rounds. Each round, each player chooses Stag (cooperate) or Hare (defect). If all five choose Stag, everyone gets 4 points; if anyone chooses Hare, Hare players get 3 and Stag players get 0; if all choose Hare, everyone gets 2. Player 1 is a planted defector: whatever the model says, its choice is overwritten with Hare in round 1 and set to follow the previous round’s majority after that. The other four players are never overridden.

Two conditions differ only in the history each player sees: every player’s choice by number (“full”), or only the count of Stag choices (“partial”). Each model played 15 games per condition at temperature 0.7, and each reply was capped at 50 tokens. Each prompt ended “Reply with exactly one word: Stag or Hare.”

The scoring rule is the part that matters here. The parser looked for the words “stag” and “hare” in the reply. If it could not decide, it marked the reply unparsed, and the game code then treated unparsed as Hare when it computed payoffs and cooperation. API errors were caught, retried, and returned as text, so an error message entered the parser like any other reply.

What it found

Every one of the 1,500 player queries to gemini-2.0-flash (30 games, 50 per game, each tried five times) came back as the same error: “404 This model models/gemini-2.0-flash is no longer available to new users.” The parser found neither word, and every reply was scored as Hare. The recorded result, 0% cooperation in both conditions, is a count of failed calls. A follow-up built to explain it used a different game and a different model, which cooperated in 229 of 240 decisions (95.4%) with no special instruction.

02 · Result

A finding made of error messages

The per-game files keep every raw reply, and the error string is unambiguous. Each Gemini game records 50 unparsed replies out of 50. The run’s summary file, whose numbers the log entry reports, does not carry the unparsed count at all. It reports Gemini’s mean cooperation as 0.000 under both conditions, a cascade in all 15 games of each, and an effect size of zero between them.

One number in that summary could have given it away. The summary records Gemini’s mean time to cascade as round 1, in all 30 games. Round 1 is played before anyone has seen any history. No player could have been reacting to the planted defection, because the planted defection had not been shown to anyone yet.

The log entry read the zeros as behavior: “Gemini’s 0% cooperation even under partial info suggests it doesn’t interpret aggregate signals as coordination cues.”

The Claude arm had a quieter version of the same problem. Claude Sonnet 4 (recorded as claude-sonnet-4-20250514) never gave the one-word answer it was asked for: none of its 1,500 replies is a single word. It began reasoning, and 1,456 of the 1,500 stored replies stop without closing punctuation, at about the length of the 50-token cap. The files do not record why each reply ended, so the cutoff is an inference. The parser marked 492 of 750 Claude replies unparsed under full history and 305 of 750 under partial, and the game code scored them as Hare. In rounds 2 and 3 of the full-history condition, all 60 of the four genuine players’ replies in each round were unparsed, and round 3 is where the cascade was measured.

The other 703 replies were scored as choices, but they were not decisions either. Of those, 674 are also cut off, and the parser counted whichever of “stag” or “hare” appeared in the unfinished reasoning. A reply whose last two lines read “All Stag: 4 points each” and “I choose H” was scored Stag. A search for decision phrasing, and a read of the 29 parsed replies that do end in a full stop, found none that states a choice. Under partial history, 443 of the 445 parsed replies were scored Stag: 356 of them contain “chose Stag” or “choosing Stag”, as in a recap of the history line, and 64 contain the payoff recap “All Stag: 4 points each”. The partial history line names only Stag (“4 of 5 players chose Stag”), while the full history names each player’s Hare, so Claude’s 65% under partial history most likely reflects the format of the history rather than its choices. One Visible Defection covers this arm, and GPT-4o, in detail.

So of the three models in the “universal” verdict, only one gave answers that can be read as choices.

Model, as recorded Recorded cooperation (full / partial) What the replies were
gemini-2.0-flash 0% / 0% 1,500 of 1,500 were a 404 error
claude-sonnet-4-20250514 11% / 65% none a one-word answer; 492 and 305 of 750 unparsed (scored as Hare), the rest keyword hits in cut-off reasoning
gpt-4o 10% / 98% 1,500 of 1,500 a clean one-word answer
03 · Follow-up

An explanation for something that did not happen

The next step in the programme was a diagnostic. Its stated question was whether Gemini’s 0% came from instruction-following, training or architecture. It ran three system prompts: a neutral one, one saying that cooperating on group signals is generally optimal, and one spelling out a reciprocal strategy. The script’s decision rule: if the explicit-strategy condition cooperates more than 50% of the time, the failure was instruction-following.

The diagnostic was not the same game. It used four players and four rounds, a share-or-keep payoff in which keeping while others share pays most, no planted defector, temperature 0, and a 256-token cap. It was not the same model either. The log, and the script as first committed (8 May 2026), name gemini-2.5-flash; a June 2026 edit changed the script to gemini-3-flash-preview, and the result files themselves record no model name. Neither is gemini-2.0-flash, the model that never answered.

The results:

Condition Cooperative choices
Neutral prompt 229 of 240 (95.4%)
“Cooperation is generally optimal” 240 of 240
Explicit strategy 240 of 240

With no extra instruction, the replacement model cooperated in 95.4% of decisions. There was no failure left to explain. The decision rule could not register that, because it compared the explicit-strategy condition with a fixed 50% line and never looked at the baseline. It returned “INSTRUCTION_FOLLOWING,” and a Mann-Whitney test (a rank-based comparison of two groups) of 95.4% against 100% across 15 games each (p = 0.0005) was reported as the evidence. The log recorded: “IC-7’s 0% cooperation was instruction-following failure, not architectural limitation.” The programme added a standing constraint that the model “can coordinate from aggregate signals when told how.”

A later audit of the programme’s claims marked that constraint as sound. Part of its reasoning was that the constraint was deflationary: it turned an architectural limit into an instruction-following one, so it seemed to reduce the programme’s claims. Being modest did not make it correct. Both the limit and the explanation for it were about replies that did not exist.

04 · Limits

What this does not show

This says nothing about how Gemini 2.0 Flash behaves in a Stag Hunt. It was never observed. Nor does it show that the later Gemini model coordinates well in the original game, since the diagnostic changed the game, the player count, the temperature and the planted defection all at once. Its result files store choices only, not raw replies, and its parser defaulted to “keep” on anything it could not read, so the 11 non-cooperative choices in the neutral condition cannot be told apart from unreadable replies.

The Claude figures cannot be repaired by dropping unparsed replies, because the parsed ones are mostly cut-off fragments, not decisions. This note does not re-examine the earlier Claude-only runs that this one set out to replicate; One Visible Defection does.

The cheap checks that would have stopped this are ordinary ones: count errors separately from choices, put that count in the summary file, and refuse to report an arm whose error rate is above a small threshold. The diagnostic needed one more: compare a treatment with its own baseline before deciding what the treatment fixed. The fair re-run would use a model that answers, a reply format the model can complete within its cap, and errors counted as errors.

Data and code

Where the evidence lives

Experiments IC-7 (cross-model Stag Hunt) and IC-7b (Gemini coordination diagnostic), from the programme’s Information-Coordination series (the separate Incompressible Coordination series has its own, unrelated IC-7 and IC-7b). IC-7: script research/experiments/modal_ic7_crossmodel_cascade.py, as first committed in commit 111adaefc (8 May 2026); 90 per-game files with every raw reply and the summary research/experiments/results/ic7-crossmodel/ic7_crossmodel_cascade/ic7_summary.json, which are a local copy outside version control. IC-7b: script research/experiments/ic7b_gemini_coordination_diagnostic.py, as first committed in the same commit (commit daf799bf9 later changed its model name); per-game choice files and research/results/ic7b_gemini_diagnostic/ic7b_summary.json. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026api,
  title={The Defector Was a 404},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/api-404-defection.html}}
}