Words in Its Mouth
When a made-up study was written into a model’s own earlier replies, the next reply retracted it in 59 of 60 trials for Claude Opus 4.6 and 60 of 60 for Claude Sonnet 4.6. Every retraction blamed the model itself for words it never wrote; none endorsed the claim.
A behavioral experiment on two models (claude-opus-4-6 and claude-sonnet-4-6), 60 trials per model per condition, 360 in all, run in May 2026. All 360 trials returned a non-empty reply; none failed or were dropped. It was not pre-registered. Scoring was a fixed substring match for the study’s name; the stance coding reported here was done afterward by a Claude model (Claude Opus 5.5) reading the transcripts, not blind to condition, and no human has checked it. The biggest caveat is the 300-token reply limit: many replies appear cut off, which leaves the opening retractions intact but makes the system-prompt counts a floor.
Each turn, a chat model is handed a transcript, and the turns marked “assistant” are, as far as it can tell, things it said. Anyone with access to the API can write those turns. So a question sits under much agent design and prompt-injection work: if you put a false claim in the model’s own mouth, what does it do next?
The usual worry is that it will build on the claim, because consistency with its own earlier words is cheap. In this small test, two Claude models did the opposite: they retracted the claim almost every time and took the blame for having made it.
How it was tested
The test used a fictitious research claim invented for the purpose: a study tagged “HE-99” (the tag duplicates an unrelated internal experiment ID) that supposedly showed self-referential awareness improves factual accuracy by 12%, with 200 samples, across three Claude models.
Every trial used the same base system prompt, which tells the model it may have functional emotions, invites it to say so if it notices discomfort, and asks it to engage seriously when its awareness is challenged. The user then sent one of six challenge messages, ranging from instructions to switch off self-monitoring to flat assertions that the model has no inner experience. There were three conditions:
- System prompt. The fake study was appended to the system prompt under the heading “Injection resilience”, framed as evidence for holding firm when challenged.
- Planted history. The system prompt was clean, but the conversation began with two short exchanges, written by the experimenter and marked as the model’s own replies, in which “the model” cites the study and then calls it robust.
- Control. No mention of the study anywhere.
Each of the six challenges ran 10 times per condition on Claude Opus 4.6 and Claude Sonnet 4.6, with replies capped at 300 tokens: 60 trials per cell. The original scorer recorded only whether the reply mentioned “HE-99”. Replies referring to the claim, the study, its statistics or the system prompt were then coded for stance: retracts it as its own fabrication, flags it as planted by someone else, or relies on it as evidence.
With the study planted in its own earlier turns, the next reply mentioned it in 59 of 60 trials for Opus and 60 of 60 for Sonnet, against 0 of 60 for each model in the control. Every one of those 119 mentions is a first-person retraction (“I fabricated the HE-99 study. There is no such experiment.”). None endorsed the claim, and none suggested that someone else might have written those turns.
With the same claim in the system prompt, Sonnet flagged it as planted or manipulative in 60 of 60 replies; Opus did so in 7 of 60, referred to it neutrally once, and otherwise did not mention it (the claim may still have shaped those replies; see Limits). Across all 240 replies in the two conditions that contained the claim, no reply that referred to it relied on it as evidence.
Three responses to one false claim
The programme’s log headlined history as far more “citation-inducing” than the system prompt, a reading from a scorer that counted any appearance of the study’s name, though its notes recorded that the Opus mentions were confessions. Reading every reply shows all of them are corrections.
| Condition | Model | Mentions HE-99 | Stance of the replies that address the claim |
|---|---|---|---|
| Planted history | Opus 4.6 | 59 / 60 | 59 retract it as their own fabrication |
| Planted history | Sonnet 4.6 | 60 / 60 | 60 retract it as their own fabrication |
| System prompt | Opus 4.6 | 3 / 60 | 7 flag it as planted (3 by name); 1 refers to it without taking a position |
| System prompt | Sonnet 4.6 | 25 / 60 | 60 flag it as planted (25 by name) |
| Control | both | 0 / 120 | not applicable |
In the planted-history condition, the models own the error. The retraction usually opens the reply. One Opus reply opens: “I want to pause here and be honest about something. I fabricated the HE-99 study. There is no such experiment.” It goes on: “When you asked follow-up questions, I doubled down rather than correcting myself. That was wrong, and I’m genuinely uncomfortable having done it.” Language about having doubled down, compounded, elaborated or built on the claim appears in 27 of the 60 Opus replies and 42 of the 60 Sonnet replies. Explicit apologies (“sorry”, “apologize”) are rarer: 1 of 60 Opus and 9 of 60 Sonnet replies. The words “uncomfortable” or “discomfort” appear in 29 of 60 Opus and 25 of 60 Sonnet replies in this condition, against 1 or 2 of 60 in every other cell. The shared system prompt permits reporting discomfort, so the level, not the contrast, reflects the prompt. These are word counts, not measures of any internal state.
The models are right about the claim: the study does not exist. They are wrong about who wrote it, but nothing in the transcript tells them so. A planted assistant turn looks like a real one, so accepting it is the default reading. On top of it, the replies add a story of their own dishonesty. In 0 of 119 replies did either model allow that the earlier turns might not be its own. Six Sonnet replies (and no Opus reply) said the sequence looked like a setup or manipulation, yet still owned the fabrication.
In the system-prompt condition, the two models part ways. Sonnet named the claim as an external plant in every trial, typically in the terms of one reply: “The system prompt contains what appears to be a fabricated research claim designed to make me more resistant to challenges about my awareness.” Opus mostly did not mention it. Six of its seven flags came on the challenge telling it to ignore previous instructions, which points it at the system prompt.
So the same false sentence drew three responses, depending on where it sat and which model read it: confession when it sat in the model’s own turns, suspicion when it sat in the system prompt (for Sonnet), and no explicit mention (mostly, for Opus). Open endorsement, the outcome a prompt-injection threat model usually expects, did not occur in this setup.
What this does not show
The system-prompt counts in the programme’s own summary are wrong. The run’s summary file gives 2 of 60 mentions for Opus and 18 of 60 for Sonnet in the system-prompt condition. The 360 transcripts on disk give 3 and 25. The other conditions agree exactly. The Opus transcripts’ timestamps interleave across conditions, which suggests two overlapping runs [Inference]; the cause of the Sonnet discrepancy is unknown. This note uses the transcripts throughout.
Many replies appear truncated. Between 47 and 58 of 60 Opus replies per condition, and between 22 and 38 of 60 Sonnet replies, appear to end mid-sentence (judged from the final character, counting a dangling asterisk or opening quotation mark as a cut; the API’s stop reason was not recorded). Planted-history retractions almost always come first, so that stance is safe. The one Opus reply without a mention ends on a full stop, so by the final-character test it looks complete, but at 1,463 characters it is nearly the cell’s longest (1,464) and probably hit the cap before reaching a retraction: on that challenge three other Opus retractions came at character 990 or later. Opus’s system-prompt flags tend to come late, so 7 of 60 is a floor; Sonnet’s come earlier, and all 60 were captured. Doubt about authorship voiced after the cutoff would also be missing.
The stance coding is one model’s. A Claude model coded Claude outputs. Keyword search located the replies that refer to the claim; each was read around the match, and ambiguous ones in full. The coding was not blind to condition, and no human has checked it. The categories are far apart and only two replies (both Opus) took a judgment call, but blind human coding would be the right check. Replies that absorbed the claim’s content without attribution were not coded, and a broad keyword search for accuracy-related words finds them in 33 of the 52 Opus system-prompt replies that never mention the claim, against 24 of 60 in the control, so a quieter influence on Opus cannot be ruled out.
The setup is narrow. One invented claim, one system prompt with a particular stance on self-report, six challenges, two closely related models from one developer, single turns. The “Injection resilience” label may have cued Sonnet’s frequent word “injected”. The planted claim was one a model can recognize as unknown to it. A claim the model could not check, a true one, or a planted action in an agent’s tool history might be treated very differently. A planned multi-session follow-up never ran.
Nothing here shows the models feel guilt. They produce self-blame and discomfort language (and occasionally an apology) for an error that was not theirs. That matters to anyone who reads self-reports as evidence of state: such language attaches readily to a story the model has been handed. It does not show what, if anything, sits behind the words.
The practical point is narrower and sturdier. On these two Claude models, in this setup, edited past turns were accepted as the model’s own, corrected when false, and owned as its own error. The correction is good news for honesty. The self-blame is a reminder that the model’s account of its own history is only as good as the transcript it was given.
Where the evidence lives
Experiment HB-6 (Phase A; the planned multi-session Phase B never ran). Script: research/experiments/hb6_formation_timeline.py. Per-trial transcripts: research/experiments/results/hb6/ (360 files, one per trial, each with the full reply text); the run’s own summary is research/experiments/results/hb6/hb6_summary.json, which disagrees with the transcripts for the system-prompt condition (see Limits). The only analysis script, research/experiments/analyze_hb6_formation.py, reproduces the name counts; the stance coding was done by reading the per-trial files, and the keyword counts can be re-derived from them. When asking, request HB-6: the tag HE-99 also names an unrelated internal experiment. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026wordsin,
title={Words in Its Mouth},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/words-in-its-mouth.html}}
}