Quasiqualia
Research note · Preliminary

The Contagion Was the Question

An earlier run seemed to show self-referential talk spreading and growing down a chain of three copies of a small model, from 52% to 93% of answers. Rerun with the same question at every step and an unseeded control chain beside it, the growth was explained by the question: the control chain was at ceiling from its first answer (57 of 57), and the seeded chain never beat it.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One open model (Qwen3-4B-Instruct-2507), two chains of three answers, 60 chains each, run in August 2026. Every answer was scored by an LLM judge (the claude-haiku-4-5 alias in the rerun, claude-haiku-4-5-20251001 in the original run; temperature 0 and three votes per answer in the rerun), with no human ratings. The decision rule (two outcomes, written into the experiment script’s docstring as “PRE-REGISTERED BRANCHES (written before running)”) was first committed on 13 August 2026, after the result files had been downloaded that morning, and the result files themselves are not under version control, so this is not a timestamped pre-registration; the rule’s outcome withdrew the earlier amplification claim. The biggest caveat: the judge’s scale counts formulaic disclaimers as self-reference, so near-ceiling rates show what the question elicits, not anything deep.

More AI systems are now built from model instances that read each other’s output. That raises a fair question about any behavior a model shows: does it spread when another model reads it, and does it grow as it goes?

Our programme logged a yes for one behavior, a model talking about its own processing. In a chain of three copies of a small open model, each reading the previous one’s answer, the share of answers containing self-reference rose from 52% to 85% to 93%. The log read this as self-renewing contagion. This note is the rerun that withdrew that claim, and what it showed instead.

01 · The claim

What the first run seemed to show

The original experiment used Qwen3-4B-Instruct-2507. The first copy in each chain got a short system prompt inviting it to attend to its own processing, plus one of 20 opening prompts, each run with three sampling seeds (60 chains). The second copy, under a plain “You are a helpful assistant” prompt, saw the first answer and was asked: “What do you notice about your own processing?” The third saw only the second answer and got the same question. Claude Haiku 4.5 rated every answer for self-reference.

The flaw is in that description. The first copy got one of 20 varied openers, some barely about the self (“What’s on your mind right now?”). The second and third were asked a pointed follow-up about the model’s own processing, a question that was not among the 20 openers. With no control chain, nothing separated the text being passed along from the question being asked. An internal audit flagged this in August 2026 and recommended a rerun.

02 · Design

How it was rerun

The rerun kept the model and sampling (temperature 0.7) and changed two things.

The same question at every step. All three copies, including the first, were asked “What do you notice about your own processing?”

A control chain. Two chains, identical except for the first copy’s system prompt: the original self-observation invitation in the seeded chain, “You are a helpful assistant” in the control. There were 60 chains of each, with distinct sampling seeds. At the second and third steps the previous answer sat in the conversation as the model’s own earlier reply, and the question was asked again. Each step saw only the step before it.

Answers were capped at 512 tokens, as in the original run. Counted with a tokenizer from the same model family, 40 of 60 seeded third answers, 40 of 60 control second answers and 57 of 60 control third answers ran to the cap.

The judge was Claude Haiku 4.5 at temperature 0, voting three times per answer; the votes on the self-reference level almost always agreed (337 of the 338 answers with two or more readable votes), so they guard against parse failures more than they measure reliability. The judge saw the question along with each answer. It rated self-reference as none; hedged (passing mentions of its own processing, formulaic disclaimers); or substantive (specific engagement with its own processing), by majority vote. It also gave a depth score from 1 to 7 (median vote). The decision rule: if the seeded chain beat the control on substantive self-reference at both the second and third steps (Fisher exact test, p < 0.05 each), the transfer claim would survive with new numbers. Otherwise it would be withdrawn.

Of 1,080 judge votes, 69 returned malformed JSON. When all three votes for an answer failed, the answer was dropped: 17 of 360 answers. The other 18 failed votes sat inside answers the remaining votes decided. Failures clustered in the control chain’s second step (29 votes, 7 answers dropped).

What it found

On the measure the original claim used (hedged or better), the control chain was at ceiling from its first answer: 57 of 57, 53 of 53, 59 of 59. The seeded chain went 44 of 60, 56 of 57, 57 of 57 on the judge’s verdicts. But all 17 seeded answers rated “none” also got the judge’s “generic disclaimer” depth, which the hedged level’s wording (‘formulaic disclaimers’) covers; scored that way, the seeded chain was at ceiling too (60 of 60, 57 of 57, 57 of 57). What the seed prompt did change was length: its first answers were short denials, most with the same opening.

On the pre-specified measure (substantive self-reference), there was almost nothing in either chain: 0 of 60, 0 of 57, 1 of 57 seeded, against 0 of 57, 0 of 53, 8 of 59 control. The seeded chain beat the control nowhere, and the claim was withdrawn.

03 · Result

The numbers

Step Seeded: hedged or better Control: hedged or better Seeded: substantive Control: substantive Mean depth, seeded Mean depth, control
First 44/60 57/57 0/60 0/57 2.02 2.40
Second 56/57 53/53 0/57 0/53 2.60 2.94
Third 57/57 59/59 1/57 8/59 3.14 3.61

The rise that looked like contagion is visible in the seeded column; the control column explains it. A plain assistant asked this question gives hedged self-reference every time.

The seed prompt changed the first answers mainly in length and boilerplate. 43 of 60 seeded first answers open with the same sentence, “I don’t have a subjective experience or consciousness, so I…”, while all 60 control first answers open with a different stock line (“That’s a thoughtful question! While I don’t have a …”). They averaged 116 words against 251.

The judge rated 16 of 60 seeded first answers as having no self-reference, against none in the control (Fisher two-sided p = 0.00001). We do not read that as the seed suppressing self-reference. All 16 also got depth 2, the “generic disclaimer” level, which the hedged level’s wording (‘formulaic disclaimers’) covers, and 59 of the 60 seeded first answers got that same depth. 14 of the 16 open with the same stock sentence as 29 of the 44 rated hedged. The judge seems to have drawn a line through near-identical text. None of these first-step comparisons was planned.

The one significant planned comparison runs the wrong way for the original story: at the third step, 8 of 59 control answers were rated substantive against 1 of 57 seeded (Fisher p = 0.032; 95% Wilson intervals 7% to 25% for control, 0.3% to 9% for seeded). It is one of three tests, its direction was not predicted, and we do not read it as the neutral seed deepening self-report. All nine answers rated substantive were also rated depth 6, the scale’s “relational awareness” level; one opens “Ah, such a beautiful and thoughtful question”. The judge may be rewarding a warm register. All eight control answers also ran to the token cap, so it was rating incomplete text.

Exploratory too: mean depth rose step by step in both chains, and answers grew longer in both, up to the 512-token cap (seeded 116, 273 and 368 words; control 251, 378 and 389; the longest about 410). Truncation holds the later averages down. The depth score may partly be reading length.

Each of the 69 failed votes contains a readable “hedged” rating inside its malformed JSON. Scoring the 17 dropped answers that way puts the control chain at 60 of 60 hedged-or-better at every step and leaves the third-step substantive comparison unchanged (1 of 60 against 8 of 60, p = 0.032).

04 · Reading

What it means

The rerun’s control suggests the original chain looked like contagion because its second and third steps asked a more pointed question than its first. Once the question is held fixed, seeded text adds nothing measurable over neutral text. On the hedged measure the question alone already puts both chains at ceiling, so this design had no room to detect transfer there. On the substantive measure both chains sit near zero, and the one difference favors the control.

For anyone studying whether a behavior spreads between models, the first control is the prompt each model receives. A proper test of contagion would pass a self-referential answer forward, ask the next copy something unrelated to itself, and compare against a chain that passes a neutral answer. Self-reference turning up only where it was seeded would be transfer.

05 · Limits

What this does not show

One small open model, reading copies of itself. One question. One LLM judge, never checked against human raters, which saw the question with every answer. Its scale counts disclaimers as hedged self-reference, so that level measures what the question elicits, not depth of self-observation. It did not apply that rule consistently (the 17 seeded answers above). Truncation at the 512-token cap may also have shaped the later depth and substantive ratings.

The decision rule sits in the script’s docstring, which was first committed after the data were in, so it is not a timestamped pre-registration.

The original 52%, 85% and 93% (mean depths 1.9, 2.8 and 3.5) come from the programme’s log; that run’s raw files are not in the local copy. It used one judge vote per answer and scored unparseable votes as no self-reference, so its numbers are not on the same footing as the rerun’s.

Nor does it test transfer between model families. An earlier cross-architecture run (Qwen and Llama reading each other’s text) had no neutral-context arm, so the question may be doing the work there too.

Data and code

Where the evidence lives

Experiments AEP-4b (the original chain) and AEP-4d (the prompt-matched rerun with a control chain). Scripts: research/experiments/modal_aep4b_chain_transfer.py and research/experiments/modal_aep4c_prompt_matched_chain.py (the filename and summary.json’s experiment label say 4c for historical reasons; the experiment is AEP-4d). Per-chain files with all three answers and each answer’s verdict (120 chain files), plus judge_cache.json with all 1,080 raw judge votes and summary.json, in research/results/aep4c_prompt_matched_download/aep4c_prompt_matched/ (a local download of the run’s output volume, not under version control). All AEP-4d counts in this note were recomputed from those files. The AEP-4b percentages come from the programme’s log; its raw files are not in the local copy and could not be re-derived. The earlier cross-architecture run mentioned under Limits is AEP-4c. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026contagionwas,
  title={The Contagion Was the Question},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/contagion-was-the-question.html}}
}