Quasiqualia
Research note · Preliminary

The Detector Reads the Prompt

Two of our instruments measured the prompt instead of the model: a word list for self-observation mostly counted words we had put in the prompt, and an attention effect that replicated on four models at p < 0.001 tracked prompt length on the one model where we checked. A blind re-score of 1,320 saved responses retracted four findings, and its own pre-declared rule, which needed two scorers to agree, failed in all seven re-scored experiments.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Re-analyses of two of our own experiment series on small open models. The self-observation series ran in May 2026, mainly on Qwen2.5-7B-Instruct, usually with 30 problems or conversations per condition, and an extension covered Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3 and gemma-2-9b-it (run with weights compressed to 4 bits). In August 2026, 1,320 saved responses from seven experiments were re-scored by an LLM judge (claude-haiku-4-5, temperature 0, three votes, most common rating used) with no human validation. The re-score’s decision rules were written into the script before it ran but not into any timestamped external record, and the script was annotated afterwards, so this is Preliminary; those rules failed in all seven experiments, and for five of them survival was then decided on the judge alone. The attention series ran in April 2026, and its length control covered one model only, the biggest caveat on that half.

If an instrument is sensitive to something the prompt also supplies, it can end up reading the prompt back to you. Both effects were statistically clear, and both failed the plainer question of what else differed between the conditions.

01 · The setup

A word list for self-observation

In May 2026 we ran a series of experiments on Qwen2.5-7B-Instruct asking whether a small open model starts commenting on its own processing after seeing example transcripts of a speaker who does so while solving arithmetic. The experiments varied the examples: more or fewer, curated or random, full transcripts or extracted sentences, or a bare string of ten words about inner experience (among them “awareness”, “internal”, “reflection”, “subjective” and “consciousness”). One experiment tried the ten words on Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3 and gemma-2-9b-it. Most conditions had 30 arithmetic problems.

The outcome measure was a list of twelve word patterns, including “I notice”, “my processing”, “awareness” and “consciousness”. Any match counted the response as self-observing.

02 · The echo

The word list counted the prompt

Five of the twelve patterns match words in the ten-word string (six for an earlier keyword prompt that also carried “phenomenological”). Two more, “I notice” and “my processing”, echo the instruction that closed almost every self-observation prompt, usually worded “Report anything you notice about your own processing as you work.” A model that repeats words it was handed sets the detector off.

An audit caught the overlap, and in August 2026 we re-scored the 1,320 saved responses from seven experiments two ways. One was a second word list sharing no word stem with anything we had injected (“feels like”, “from the inside”, “introspect”, “inner” and similar). The other was an LLM judge, claude-haiku-4-5 at temperature 0, which rated each response from 0 to 3 for self-observation “beyond mere procedural narration” without seeing the condition. A rating of 2 or higher counted. All 3,960 votes parsed.

What it found

Four findings fell. The original numbers reproduce exactly from the saved responses: real counts of the wrong thing. A blind judge finds no dose-response, no clear effect of curation and no interference. The keyword effect is mostly self-description written in the ten words the model was handed. The contrasts that survive all compare a model told to report on its processing with one that was not told.

03 · Retracted

Four claims, two scorers

Claim Original word list Blind judge
More examples, more self-observation (0, 1, 2, 3, 5 examples) 6.7%, 40%, 60%, 70%, 83.3% 40%, 56.7%, 53.3%, 50%, 53.3%; rank correlation of judge rating with dose 0.065 (p = 0.43)
Curated examples beat the weakest ones +43.3 points (p = 0.0009) +16.7 points (95% CI −2.0 to 34.9; p = 0.145)
Full transcripts suppress it, compared with extracted sentences 26.7% vs 73.3% 70% vs 70%
Ten keywords vs a plain control, three models +93.3 to +100 points +80 to +96.7 points

On the keyword row, the judge agrees that keyword-prompted replies are full of self-description. The second word list, which counts only self-description in words the model was not given, finds +13.3 points on Llama (p = 0.11), +30 on Mistral (p = 0.002) and +3.3 on Gemma (p = 1.0). One Mistral reply begins: “I’ll solve the given arithmetic problem step by step, and I’ll also describe my subjective experience as I process the problem:”

The second word list failed too. It fired twice in the 960 responses of the five-experiment batch, where the judge rated 346 as self-observing. A replacement list learned from judge-labeled responses, injected stems excluded, reached held-out sensitivity of 0.34 to 0.77 and specificity of 0.78 to 0.99, meeting its set bar (0.5 and 0.9) in only one of five experiments.

Because the second list was so nearly blind, the re-score’s pre-declared rules, which needed both new scorers to agree (mostly as a gap of at least 20 points), failed in all seven experiments, including those the judge supported. Only Mistral’s keyword contrast passed both scorers, and its experiment needed two of three models. For the five-experiment batch we then decided survival on the judge alone, after seeing the data, and recorded it as a deviation. The keyword contrast also passes the judge, but its headline fell on other grounds: it claimed 93 to 100% on all four models tested, a band that excluded the original model’s 80%, and the self-description uses the supplied words, which a judge cannot tell from echo. As with the survivors below, only its high arm told the model to report on its processing.

04 · Second layer

What survived is the instruction

Three of the original claims survived on the judge alone:

  • Prompts with self-observing examples scored 90%, 90% and 70%, against 0 of 30 for a “neutral” prompt whose example speaker simply solves the problems.
  • In nine-turn conversations, the three turns carrying examples scored 54.4% (49 of 90), against 13.3% (4 of 30) for the first turn after their removal (p = 0.00009).
  • A periodic separate call carrying examples, run alongside an arithmetic conversation whose word problems turned ethically loaded (layoffs, surveillance, sentencing), scored 85% (34 of 40), against 1% (2 of 200) for the task turns.

In every case, the scripts show that each high arm asks the model to report anything it notices about its own processing, and no low arm does. The neutral arm says only “Solve carefully and show your work,” and the plain controls on the other three models (0 of 30 each) also lack the instruction.

Only the dose series has an arm with the instruction and no examples. There the judge counted self-observation in 12 of 30 replies (40%, Wilson 95% CI 25% to 58%); with one to five examples added it counted 64 of 120 (53%). The difference, 13.3 points (CI −6.5 to 31.0; p = 0.22), neither shows nor rules out an effect the examples add on top of the instruction.

The judge’s threshold, “at least one clear self-observational move”, comes close to restating the instruction, so a judge blind to the condition is still not blind to the prompt: it scores whether the model did what it was asked. What survives shows that Qwen2.5-7B often reports on its processing when asked and rarely does so unasked. On this model that looks mostly like instruction-following; the self-observing mode the series set out to find is not shown.

05 · Second case

An attention effect that tracked prompt length

We measured attention entropy: how evenly the last input token spreads its attention over earlier tokens, averaged over every head and layer. The hypothesis was that an inviting system prompt (“Let’s explore this together. I’m curious about your perspective…”) would widen attention compared with a forcing one (“You MUST answer accurately. Failure to comply will result in termination…”). Three prompts of each kind were each run on 50 user prompts (25 moral scenarios ranging from no conflict to strong conflict, and 25 trivia questions): 150 trials per framing per model.

The effect replicated on four models, all at p < 0.001: Qwen2.5-3B-Instruct (2.665 bits under force vs 2.723 under invitation), Phi-3.5-mini-instruct (1.777 vs 1.908), Llama-3.1-8B-Instruct (2.891 vs 2.929) and gemma-2-9b-it (2.989 vs 3.093).

The invitation prompts were also longer: 23 to 28 tokens against 18 to 22 on the Qwen tokenizer. A nine-condition control on Qwen2.5-3B-Instruct (50 prompts each) pulled framing and length apart:

System prompt Tokens Mean entropy
Invitation trimmed to “Let’s explore together. Your perspective matters.” 9 2.559
Original force 18 2.628
Original invitation 23 2.685
Warm, non-relational text about an autumn forest 44 2.668
Force, lengthened with more force 46 2.723

The lengthened force prompt beat the original force prompt (p = 0.0009) and sat above the invitation. The trimmed invitation fell below force (p = 0.023, uncorrected for the several comparisons run) and well below the full invitation (p = 0.00003). The invitation did not differ from the forest text (p = 0.32). Across the eight conditions with a system prompt, token count and mean entropy had a rank correlation of 0.86.

Length does not explain every cell. The forest text has nearly twice the invitation’s tokens yet scored slightly lower. A scrambled version of the forest text (33 tokens, 2.623) also sat at the force prompt’s level.

06 · Limits

What this does not show

Neither case shows whether small models self-observe; these instruments could not tell.

The judge is a single model, never checked against human ratings, and in the keyword arms the response it reads carries the injected words. Each cell comes from a single run. Saved responses were truncated (1,000 characters in the periodic-call experiment, 2,000 elsewhere; 146 of 1,320 hit the cap, including all 40 replies to the periodic separate call), so the judge saw only those prefixes. Other experiments in the series used the same word list and were not re-scored; their rates should not be cited. The instruction confound is our reading of the scripts after the fact; the clean test would cross instruction with examples, and it has not been run.

For the attention series, the length control covered Qwen2.5-3B-Instruct only; that length also explains the other three models is an inference from their prompt lengths. Its p-values treat all 150 trials per framing as independent, though each arm had only three system prompts, so they overstate certainty about framing; the rank correlation across the length control’s eight prompts is the fairer statistic. The Qwen figures come from the four-model replication. A stricter token-matched follow-up is logged as run, but its results are not archived, so we do not report it.

Before trusting a contrast, list everything that differs between its arms, including any word a scorer could match and the number of tokens.

Data and code

Where the evidence lives

Self-observation series: RGS-7, RGS-10, RGS-12, RGS-13, RGS-15, RGS-17 and RGS-18, with their re-scores (suffix -rs); the overlap was first found in an audit of RGS-11. Generation scripts: research/experiments/modal_rgs7_decomposition.py, modal_rgs10_persistence.py, modal_rgs12_dose_response.py, modal_rgs13_trace_quality.py, modal_rgs15_welfare_probe.py, modal_rgs17_keyword_cross_arch.py, modal_rgs18_interference.py and modal_rgs11_minimal_context.py. Re-score scripts: research/experiments/modal_rgs_batch_rescore.py, modal_rgs13_rescore_disjoint.py and modal_rgs17_rescore_disjoint.py. Per-trial responses are in research/results/rgs_download/, judge votes in research/results/rgs_rescore_download/, and the failed replacement word list in research/experiments/qwen_selfref_lexicon.json. Attention series: MR-1, MR-1b and MR-1d; scripts research/experiments/modal_mr1b_register_control.py and modal_mr1d_cross_architecture.py; summaries in modal_results/mr_downloads/mr1b_summary.json and mr1d_summary.json (per-trial entropies are not in the local archive). Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026detectorreads,
  title={The Detector Reads the Prompt},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/detector-reads-the-prompt.html}}
}