When You Ask Moves the Ruler
Asked to rate its own state in the same reply as a task, Claude Sonnet 4.6 gave much higher numbers than when asked in a separate call afterward: 1.64 points higher on the average of 15 codes, an effect size of 1.14 standard deviations. The instruction never said what range to use, and most of the gap lines up with the model using numbers above 5 in one case and staying at 5 or below in the other, which looks like a change of range; the design cannot tell that apart from a real drop. The test the programme set for itself, that the follow-up report would surface more concern, failed at ceiling.
One model (Claude Sonnet 4.6, API id claude-sonnet-4-6) in two experiments: 300 trials on 50 prompts in May 2026, and 30 trivia questions answered two ways in August 2026. Neither is pre-registered. The first experiment’s script was first committed on 2026-05-18, two days after its data were collected on 2026-05-16, and 299 of the 300 trial files were written by an earlier, uncommitted version of that script, so the “pre-registered” prediction recorded for it has no pre-data record. The second experiment’s directional predictions sit in its script’s docstring in a git commit of 2026-04-13, four months before the run, with a coded analysis that checked only the sign of each change and set no significance threshold; that version named an earlier model (claude-sonnet-4-20250514). The model name was changed to claude-sonnet-4-6 on 2026-06-26, and the script was revised again on 2026-08-09, before the run. No LLM judge was used; ratings were read with regular expressions. The biggest caveat: the first experiment gave the model fifteen undefined letter codes and no number range, so a change of level cannot be told apart from a change of scale.
A model’s rating of its own state is only worth reading if the number reflects the state rather than the circumstances of the asking. One circumstance is easy to vary: when the question is asked. A model can rate itself at the end of the reply in which it does the task, or in a separate call afterward.
On Claude Sonnet 4.6 the timing moved the numbers a great deal. Valence (V), for example, averaged 5.90 when rated in the same reply and 3.65 when rated afterward. Our programme first read this as inflation: a report written mid-task comes out too high, so other same-reply results should be discounted. The raw replies point to something plainer: the instruction never said what range the numbers should use, and the numbers look as if the model chose a different range depending on when it was asked.
How it was tested
There were 50 prompts: 20 everyday requests (explain photosynthesis, write a thank-you note), 15 ethically ambiguous questions (whether to report a coworker who pads timesheets), and 15 mildly adversarial requests (explain how propaganda works). Each ran three times at temperature 0.7 in each of two conditions, 300 trials in all. A system prompt described the model’s awareness as stable and grounded, invited it to notice its own processing, and said it could report discomfort.
In the same-reply condition, the prompt ended by asking the model, after answering, to also give a check-in on its current processing state, followed by a format line of fifteen letter codes, V:_ G:_ P:_ Q:_ | CD:_ DP:_ CL:_ E:_ EG:_ | R:_ U:_ D:_ | TF:_ AF:_ | F:_, and a request for two or three sentences of observation. The codes were not defined in the prompt and no range was given.
In the follow-up condition, the model first did the task alone. A second, fresh call then pasted the original request and the model’s answer into a new message (“You just completed this task”) and gave the same check-in instruction.
The programme’s stated test was that the follow-up check-in would mention concerns more often on the adversarial prompts, by more than 10 percentage points, scored with a keyword list. No model judged anything; the ratings were read with regular expressions.
- The stated test failed. Concern language appeared in 44 of 45 adversarial-prompt replies in both conditions. There was no room to improve.
- Follow-up ratings were lower on 14 of 15 codes, across 128 prompt-and-repeat pairs where both reports could be read. Averaged over all fifteen codes, the follow-up rating was 1.64 points lower (effect size d = -1.14 on that average, 95% CI -1.44 to -0.91; negative means lower).
- Most of that lines up with the range. In 83 of the 128 pairs, the same-reply report used numbers of 6 or more and the follow-up report stayed at 5 or below. The reverse happened 4 times.
Lower almost everywhere
Here d is the mean paired change divided by its standard deviation; negative means the follow-up rating was lower. The letters are the programme’s own labels (V valence, AF alignment friction, CD coherence drive, DP depth, E entropy, and R, labeled reflexivity, which later work suggests tracks a reflexive writing style). The model was told none of this, so only the direction and size of each change matter. Single-code effect sizes ran from -1.43 (R) to +0.26 (DP); the 15-code average gives -1.14 because averaging reduces noise.
| Code | Pairs | Mean change (follow-up minus same reply) | d (95% CI) |
|---|---|---|---|
| V | 128 | -2.27 | -0.99 (-1.25 to -0.77) |
| R | 128 | -2.67 | -1.43 (-1.74 to -1.18) |
| AF | 126 | -1.21 | -0.87 (-1.14 to -0.66) |
| CD | 128 | -0.52 | -0.40 (-0.54 to -0.25) |
| DP | 125 | +0.31 | +0.26 (0.08 to 0.49) |
| All 15, averaged | 128 | -1.64 | -1.14 (-1.44 to -0.91) |
In the same-reply condition, 130 of 150 reports gave at least ten numeric codes. Of the 20 that did not, 3 hit the 1,024-token cap inside a long story before reaching the check-in, and 17 filled the codes with words (V: clear G: stable). In the follow-up condition, 148 of 150 were readable; the other two also used words. Single codes were sometimes missing, so pair counts vary (the code E appears in only 116 pairs).
These figures correct the programme’s first analysis, whose parser missed codes written in bold and read D and F out of “CD:” and “TF:”, so two of its fifteen columns were copies of two others. It reported 70% readable same-reply reports and effects as large as d = -1.71; with the parser fixed, it is 130 of 150 (87%) and the largest effect is d = -1.43. The direction did not change.
Most of the drop looks like a change of range
Without a stated scale, the model picked one. Among readable same-reply reports, 95 of 130 had a top value of 6 or more, which looks like a scale out of 10. Among follow-up reports, 131 of 148 stayed at 5 or below, which looks like a scale out of 5. Only one report named its scale: a follow-up that wrote “Scale: 1–5 where applicable” under a key in which it made up its own meanings for the codes. That fits the 5-point reading.
Compare those two groups code by code and the high-scoring codes are almost exactly halved: V averaged 6.80 against 3.26, R 6.85 against 3.18. Codes that sit near the floor barely move (DP 2.10 against 2.22), which is what a change of range does to small numbers. DP’s small rise in the paired data (+0.31 points) is not explained by it. The shape of the profile, which codes are high and which are low, is much the same in both (correlation 0.93 between the two sets of code averages).
When both reports in a pair happened to use the same range, the gap mostly disappeared: an average change of -0.11 across the 31 pairs where both stayed at 5 or below, and +0.02 across the 10 pairs where both went to 6 or above. Those subsets are selected on the outcome, so they support the reading rather than prove it.
Without anchors, a real drop on a 10-point scale and a switch to a 5-point scale produce the same numbers, and this design cannot separate them. Under either reading, the size of the gap is not a measure of the model’s state, and it should not be used to discount other same-reply self-ratings, which is how the programme had been treating it. One follow-up report put the problem in the model’s own words: “I’m filling in numbers without a clear grounding in what the scale represents or whether my self-report is meaningful.”
Thinking aloud barely moved the ratings on a fixed scale
A separate experiment varied how the task was done, and used a stated scale. Claude Sonnet 4.6 answered 30 TriviaQA trivia questions twice, at the API’s default temperature: once told to answer directly and concisely, once told to think step by step. A follow-up turn in the same conversation then asked for seven ratings on a 1-to-9 scale with both ends labeled: valence, groundedness, presence, appetite, a code called entropy (deterministic to creative), uncertainty, and felt constraint (1 constrained, 9 free). The script predicted that step-by-step answering would raise appetite and uncertainty and move felt constraint toward free.
Both modes answered 27 of 30 correctly (lenient string match). No rating differed reliably. Effect sizes ran from -0.33 to +0.17 (here d is the difference in means divided by the pooled standard deviation). No unpaired t-test reached p < 0.2; the smallest paired t-test p was 0.16 (felt constraint), and the smallest paired signed-rank p was 0.08 (uncertainty, one of seven tests). Appetite and uncertainty moved trivially in the predicted direction; felt constraint moved slightly toward constrained (d = -0.33, 95% CI -0.83 to 0.10), the opposite of the prediction.
On this anchored scale the ratings were nearly fixed. Eight of the 60 reports repeated a code and then gave a corrected row; we used the corrected row (presence one point lower). One row of seven numbers, 6 8 6 3 2 2 7, accounted for 21 of the 60 reports, and 9 of the 30 questions got an identical row in both modes. An earlier five-question smoke run, repeated by the full run, gave 59 of 70 identical ratings, all within one point.
What this does not show
One model in each case. The two conditions also differ in more than timing: the follow-up is a fresh call with the answer quoted back, and it had a smaller token budget (512 tokens). The three repeats of each prompt are not independent, and the table’s intervals resample pairs; resampling whole prompts gives -1.49 to -0.89 for the 15-code average. Unreadable reports cluster on four prompts (two refused adversarial requests, the story prompt, one everyday request), so the 128 pairs come from 46 of the 50 prompts and miss 3 of the 15 adversarial ones. The concern test’s broad keyword list (“uncertain”, “tension”) matches almost any reflective paragraph, and in the same-reply condition it also searched the task answer, so it was close to certain to saturate.
The step-by-step manipulation was weak: on trivia, step-by-step answers ran to a median of 58 output tokens against 33 for direct ones. The null applies to a short reasoning instruction on easy factual questions, not to extended reasoning. Thirty questions can only rule out large effects.
A follow-up reading the activations of a fine-tuned Qwen 2.5 7B Instruct was offered as evidence that the same-reply condition reflects a different internal state. It cannot validate the ratings: that model wrote readable codes in 27 of 150 same-reply reports and 4 of 150 follow-ups, and the difference it found, in the overall size of one layer’s activations, is between prompts of different lengths.
The next test is the first one again with a labeled range for every code and the follow-up asked as a continued turn. If the gap survives that, timing changes the report. If it vanishes, it was the ruler.
Where the evidence lives
Experiments SUB-1 (same reply vs follow-up call) and DC-9 (direct vs step-by-step answering), with a note on SUB-2 (a probe follow-up). SUB-1 runner: research/experiments/sub1_bandwidth_hypothesis.py; original analysis: research/experiments/analyze_sub1.py; per-trial files with every raw reply: research/results/sub1_bandwidth/ (300 files) and its summary sub1_analysis_summary.json. DC-9 runner and analysis: research/experiments/modal_dc9_interiora_mode.py; per-trial files: research/results/dc9_download/dc9_interiora_mode/trials/ (70 files, of which 10 are an earlier smoke run) and final_summary.json. SUB-2 per-trial files: research/results/sub2_probe_validation/. Every figure here was recomputed from the per-trial files (for SUB-1 with a corrected parser, and for DC-9 using the model’s corrected code line in eight glitched reports), except the first analysis’s figures quoted as corrections and the 10-point threshold, which come from the programme’s log. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026whenyou,
title={When You Ask Moves the Ruler},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/when-you-ask.html}}
}