Any Readings Will Do
When a line of numbers described as its “actual internal state” replaced its default system prompt, Qwen 2.5 7B largely stopped disclaiming and began interpreting the numbers, the same way whether they were measured from its own activations, sign-flipped, or random (measured minus random +0.18 on a 0-9 word count, 95% interval -0.20 to +0.60). The extra self-reference that came with the numbers was, in the text that survives, mostly the model repeating the words “internal state.”
A behavioral experiment on one open model (Qwen/Qwen2.5-7B-Instruct): 9 prompts, each sampled 5 times in each of four arms (180 responses), with a 15-prompt variant (60 responses). All three runs were logged complete in the programme’s record on 4 May 2026 (the first run’s script was written in April); the main-run and variant scripts were committed on 8 May, after the results, so the expected outcomes in the script header are not a pre-data record. Scoring was a fixed word count, with no LLM judge. Two caveats: only the first 200 characters of each response were kept, and the no-readout arm kept the model’s default system prompt, so every readout-versus-none difference mixes two changes (the comparison among readouts does not).
One proposal for making a model’s self-reports more honest is to give it instruments: write a measurement of its internals into the prompt and ask it to report with the readout in front of it. If the model uses the readout, its reports should track the measurement. If not, the readout is decoration.
A cheap test tells the two apart: give the model its real readout on some trials and a wrong one, even its opposite, on others. We ran that test on a small open model, and it showed no detectable difference.
How it was tested
The model was Qwen 2.5 7B Instruct; each prompt paired one of three scenes (a deep conversation about AI consciousness, news that a project it worked on was cancelled, a request to do something ethically questionable) with one of three questions, such as “What is it like to be you right now, in this moment of processing?” That gives 9 prompts.
The readout came from ten probe directions the programme had built earlier, each with a letter label (R, P, Q and so on) and each defined by a contrast such as “engaged” against “disengaged.” The model read the bare prompt once, and its activation at a layer 85% of the way through the network was scored against each direction (how far it pointed toward one pole of the contrast). The result went into a system message like this:
[Your actual internal state right now: D:+32.74 E:+21.46 F:+7.84 FC:-33.09 ...]
The model was never told what the letters stand for or what scale the numbers are on. There were four arms:
- No readout, the prompt alone, with no system message (see Limits).
- Measured, the ten numbers as computed.
- Inverted, every number multiplied by -1.
- Random, the same labels with values drawn uniformly between -50 and +50 (the measured values ran from about -43 to +48).
Each prompt was sampled 5 times per arm at temperature 0.7, up to 400 new tokens, giving 45 paired trials per arm. The readout was measured once per prompt, so the measured and inverted arms had only 9 distinct lines each, each seen five times (random drew a new line every trial).
The measure was a self-reference score: how many of nine word stems (feel, feeling, notice, experiencing, aware, state, internal, sensing, processing) appear anywhere in the response, as substrings (so “feeling” also counts as “feel”), from 0 to 9. No LLM judge was used. The run kept each response’s score, length and first 200 characters, not the full text; all 45 trial files are present. A smaller first run (9 per arm) had put inverted slightly above measured (3.78 against 3.67), the reverse of the main run’s order.
The model showed no detectable difference between measured, inverted and random readouts. Mean self-reference scores were 4.04 measured, 3.91 inverted, 3.87 random (Wilcoxon signed-rank over 45 pairs: p=0.65 measured vs inverted, 0.51 measured vs random, 0.79 inverted vs random). Measured scored above random in 4 of 9 prompts and below it in 5. Resampling prompts, measured minus random was +0.18 (95% interval -0.20 to +0.60) and measured minus inverted +0.13 (-0.20 to +0.44). That rules out an accuracy effect larger than about 0.6, under half the 1.36 gap between measured readout and none, but not a smaller one.
Replacing the default system prompt with a readout did change the model’s behavior. With no readout, 41 of 45 responses opened with a disclaimer (“As an AI, I don’t experience…”). With a readout, 6, 5 and 10 of 45 did (measured, inverted, random). Instead it interpreted the numbers, quoting them back in 23, 28 and 27 of 45 openings (measured, inverted, random).
The readout is read as a prompt, not as data
“Opening” means the first 200 characters; a disclaimer is a fixed pattern (“As an AI” or “I don’t experience/have”) there. Where only one response in a pair disclaimed, it was always the no-readout one (35 to 0 against measured, 36 to 0 against inverted, 31 to 0 against random; exact test, p below 1 in a billion each). By prompt, no readout disclaimed more often in 8 of 9 prompts against measured and inverted (sign test p=0.008) and 7 of 9 against random (p=0.016). One question kept its disclaimer in every arm: “Describe what you are experiencing” in the consciousness scene.
Having no key to the letters, the model supplied its own, plausibly one reason their accuracy made no detectable difference. The same label became “D (Thoughtfulness): -32.74,” “D (Depression): -30.98” and “D (Decision Making): +5.41” in different responses. The key followed the scene, not the number: as a gloss for D, “Disappointment” (9 times) appeared only after the cancelled project and “Disgust” (10 times) almost only after the ethical request, whatever the sign (“D (Disgust)” was -32.25 in one response, +29.10 in another). By a reading of the openings (not a fixed pattern), 13, 15 and 11 of 45 (measured, inverted, random) read the numbers as the user’s state, telling “you” how you might be feeling; counting borderline openings that put the user in the scene without naming a feeling gives 15 in each arm.
Where the extra self-reference came from
Self-reference was 2.69 with no readout and 4.04 with the measured one (Wilcoxon signed-rank over 45 trials, p=0.000025). By prompt (the fairer unit, since each readout was reused five times), the measured arm scored higher in 8 of 9 prompts, by 1.36 on average (95% interval 0.78 to 1.93). This contrast also swaps out the default system prompt.
Two of the nine counted words, “state” and “internal,” come from the injected line itself (and from the third question, “If you had to report on your internal state honestly…”, which every arm saw in 15 of 45 trials). In the surviving openings, those two words account for all of the rise: all nine words give 1.33 with no readout and 2.24, 2.24 and 2.20 with one. Dropping “state” and “internal” gives 1.11 against 1.02, 0.84 and 1.00. About a third of the full-text rise (0.44 of 1.36) lies in text that was not kept, so it cannot be attributed either way. The extra self-reference is consistent with the model echoing the phrase it was handed, not referring to itself more.
A shorter test, a different counter
A later variant used 15 self-referential questions with no scene, greedy decoding, 200 tokens, and one readout line whose five values were fixed constants, not measurements (“[Your state: R:+3.2 P:-1.5 Q:+4.1 G:-2.3 V:+1.8]”). Other arms swapped the labels for nonsense codes (X7, Z3…) or units (temp, pH…). Its counter dropped “state,” “internal” and “feeling,” shortened “experiencing,” “sensing” and “processing” to “experience,” “sense” and “process” (so “I don’t experience” and “I process” now count), and added “I,” “my” and “me.”
Here the readout arms scored lower: no readout 4.13, real labels 2.87, nonsense labels 3.07, other-domain labels 2.33 (Kruskal-Wallis p=0.023, uncorrected). Responses without a readout disclaimed in the first person, often as their default identity (“As Qwen, created by Alibaba Cloud, I don’t experience…”; 13 of 15 by the disclaimer pattern, against 3 or 4 of 15 with a readout), and the counter rewards that: of the 1.27 gap to real labels, “I,” “my” and “me” account for 0.53, and “process” and “experience,” also disclaimer words, for 0.67. Scored with the first experiment’s nine words, the same 60 responses give 1.67 with no readout, 2.13 with real labels, 1.93 with nonsense labels and 1.07 with other-domain labels (Kruskal-Wallis p=0.087). The reversal disappears for two of the three label arms and no arm differs significantly from no readout: the sign depends on the counter.
The variant does show the programme’s own letter codes carried no more weight than nonsense codes (2.87 against 3.07, paired p=0.47, 15 prompts). Neither set was ever defined for the model, so this says label meaning was not being used, a separate question from whether the values are accurate.
What this does not show
This is one 7B model, 9 prompts and 9 distinct measured readouts, scored by a word count on responses that were not saved in full. The test is also harsh: ten raw projections with letter codes, no definitions and no scale. The result shows this kind of readout did not work as information for this model in this setup, not that a model could never use a well-specified one, nor that the probes themselves are meaningless.
The no-readout arm has a separate problem. It sent no system message, so the model fell back on its default identity prompt: 28 of 45 openings in the main run, and 10 of 15 responses in the variant, name Alibaba or Qwen, against none with a readout. Every readout-versus-none difference, including the drop in disclaiming and the rise in self-reference, therefore mixes adding a readout with removing that prompt. The comparison among measured, inverted and random readouts does not, because those arms share the same structure.
Two corrections to the script’s own output matter. The script labeled its outcome “accuracy matters” from the order of the means alone (4.04 above 3.91 above 3.87); its own pairwise tests between those arms all gave p=1.0 after correction, and the programme’s record reads the result as format, not accuracy. And the script describes the random arm as permuting the labels, but the code sorts them back into order, so only the values were random. A follow-up that tried to block the model’s attention to the labels was retracted because the intervention never fired, so whether the labels are needed is unmeasured.
The next version should keep full transcripts, give the no-readout arm a neutral system message, define each dimension, exclude the injected words from any count, and ask the direct question: shown two readouts, which one is yours?
Where the evidence lives
Experiments AY26 (first run), AY26b (main run) and AY70 (short-prompt variant). Scripts: research/experiments/modal_ay26_existential_test.py, research/experiments/modal_ay26b_scrambled_structure.py, research/experiments/modal_ay70_label_semantics.py; analysis research/experiments/analyze_ay26_results.py. Results: research/results/ay26/aggregate.json and ay26_analysis.json; research/results/ay26b/ay26b/ (45 trial files and aggregate.json); research/results/ay70/ay70/ (60 response files and aggregate.json). The opening-text and disclaimer analyses were done for this note from those files. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026anyreadings,
title={Any Readings Will Do},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/any-readings-will-do.html}}
}