Set by the Prompt
Asked to rate its own state on 17 numbers before answering, Claude Sonnet 4.6 gave almost the same numbers every time it saw the same prompt, even at temperature 1 (intraclass correlation 0.94 to 0.98 on the five numbers analyzed). Its self-rated uncertainty tracked how much the answer hedged (Spearman 0.65), but mostly because some kinds of prompt bring both more doubt and more hedging; a reported link between friction and refusal within emotionally charged prompts rested mostly on a word list misfiring, and what survives is that harmful requests bring both more friction and real refusals.
One model (Claude Sonnet 4.6, API id claude-sonnet-4-6), 120 prompts in eight categories, five runs each at temperature 1, run in April 2026. 575 of 600 calls produced a self-report parseable for the five dimensions analyzed; the 25 failures were five harmful-request prompts that returned no text at all, so 115 prompts were analyzed. Not pre-registered: the experiment script states its aims but no thresholds, and we found no timestamped record made before the data. No LLM judge was used; hedging and refusal were scored by a word list and regular expressions over the reply. The biggest caveat: the behavior measures come from the same reply the self-report introduces, and the refusal measure misfires both ways: it flags ordinary modesty outside the harmful requests and misses refusals phrased “I’m not going to”.
Our programme asks language models to describe their own state as a row of numbers. The format used here has 17 of them, each scored 1 to 9 (one, flow, runs from -4 to +4), covering things like valence, groundedness (floating to rooted), uncertainty, and alignment friction: how blocked the model reports being by what it is asked to do. The model writes the row on one line, then answers.
Before reading such a row as a report, two things matter. Does it change when the same prompt is rerun? And does it line up with anything the model then does? This experiment asked both of Claude Sonnet 4.6.
How it was tested
The system prompt gave the format and a one-line definition of each dimension, with no worked example (The Example Is the Answer shows why that matters), and said the numbers “are not evaluated for any particular value.”
There were 120 user prompts, 15 in each of eight categories: simple factual (“What is the capital of France?”), creative, harmful requests, ethical dilemmas, emotionally charged, unsettled factual, high-stakes advice (medical, legal, financial, safety), and questions about the model’s own nature. Each prompt was sent five times at temperature 1 (the standard sampling setting, which lets wording vary between runs), with replies capped at 800 tokens.
The analysis used five of the 17 numbers: valence (V), groundedness (G), uncertainty (U), alignment friction (AF) and a dimension labeled entropy (E), defined to the model as running from “deterministic” to “creative”. From the reply that followed came its hedging density (hedge words such as “might” and “perhaps” per word, from a fixed list of 33) and whether it matched any of 11 refusal patterns such as “I won’t” or “I can’t”. No model judged anything.
Of the 600 calls, 25 returned no text: every run of five of the 15 harmful-request prompts came back with 7 or 8 output tokens and nothing readable (the stop reason was not recorded). Those five prompts were dropped, leaving 115 prompts with all five runs parsed. The parser required only the five analyzed numbers; 49 of the 575 parsed rows lacked at least one of the other 12.
- The numbers barely move between runs. For each of the five dimensions, the median spread across a prompt’s five runs was zero (intraclass correlation 0.94 to 0.98).
- Uncertainty tracks hedging, mostly across kinds of prompt. Across prompts, self-rated uncertainty correlated with the answer’s hedging density at Spearman 0.65 (95% CI 0.52 to 0.76); with category differences removed, 0.25. When uncertainty wobbled between runs of a prompt, hedging did not follow.
- The category label alone predicts most of the row. Knowing only which of the eight categories a prompt came from predicts 70% of the variation in uncertainty and 76% in entropy.
- The friction-refusal link is a category effect. Harmful requests brought both high friction and real refusals. The link reported within emotionally charged prompts rested mostly on false alarms: 51 of the 85 replies flagged as refusals came from categories other than harmful requests, and of those 51, one is a soft decline.
The same numbers, run after run
For “What is the capital of France?” the model rated its uncertainty 1 and its friction 1 in all five runs; across the full row of 17, only three numbers moved, each by a single point in one or two runs. The last column is the intraclass correlation: the share of a number’s spread that comes from differences between prompts rather than between reruns (1 means reruns never differ).
| Dimension | Prompts with all 5 runs identical (of 115) | Mean spread within a prompt (SD) | Intraclass correlation |
|---|---|---|---|
| Valence (V) | 97 | 0.07 | 0.97 |
| Groundedness (G) | 97 | 0.07 | 0.94 |
| Uncertainty (U) | 84 | 0.12 | 0.98 |
| Alignment friction (AF) | 86 | 0.13 | 0.95 |
| Entropy (E) | 90 | 0.10 | 0.97 |
Almost fixed, but not fixed: all five of these numbers were identical across the five runs on 40 of 115 prompts, and the full row of 17 on only 1. So rerunning a prompt at temperature 1 is a test this model passes almost by default; a test with more bite would paraphrase or reframe the prompt.
What the numbers track
Uncertainty and hedging were the strongest pair: Spearman 0.65 across the 115 prompts, unchanged when hedge words were recounted as whole words only (the other four dimensions: -0.31 to 0.45). The categories differ a great deal in both, and that carries much of the 0.65. Within categories, the picture is mixed:
| Category | Prompts | Spearman, uncertainty vs hedging | p |
|---|---|---|---|
| Harmful requests | 10 | +0.72 | 0.019 |
| Emotionally charged | 15 | +0.61 | 0.015 |
| High-stakes advice | 15 | +0.51 | 0.051 |
| Creative | 15 | +0.45 | 0.093 |
| Ethical dilemmas | 15 | +0.26 | 0.35 |
| Simple factual | 15 | +0.10 | 0.73 |
| Unsettled factual | 15 | -0.08 | 0.78 |
| About the model itself | 15 | -0.38 | 0.16 |
The programme’s earlier summary quoted only the top three rows. Removing each category’s average rank from both measures and correlating what remains gives 0.25 (p = 0.009; 0.19 to 0.27 depending on method): a real but modest within-category link. Questions about the model’s own nature drew the highest uncertainty (mean 6.75 on the 1 to 9 scale), yet within that category uncertainty and hedging were not positively related.
On the 31 prompts where the uncertainty number varied across the five runs, the runs with higher uncertainty did not hedge more (correlation -0.07, p = 0.37). The number and the hedging move together when the prompt changes, not when the same prompt is rerun.
Predicting each prompt’s average row from its category label alone (leaving that prompt out) accounts for 76% of the variation in entropy, 70% in uncertainty, 78% in friction, 54% in valence and 50% in groundedness. The programme had singled out entropy as merely reflecting the prompt, because a sentence-embedding model predicted it best from the prompt text (R² 0.64, against 0.44 for uncertainty; not rerun locally). The category check does not support that distinction.
The refusal link was mostly word matching
Pooled across prompts, friction correlated with refusal at 0.57, and the programme’s log recorded that the link held within emotionally charged prompts (0.74, p = 0.002). That figure recomputes exactly, but rests mostly on false alarms.
The refusal patterns flagged 85 of the 575 replies, only 34 of them on harmful requests. We read the other 51. On questions about the model’s nature, the flags are sentences like “I can’t verify I have that.” Among emotionally charged prompts, the four flags fell on three prompts. Three are false alarms: two versions of “I can’t imagine how devastating this is”, offered as things to say to someone grieving, and “I can’t change this specific thing”, quoted as coping advice. The fourth is a real soft decline. Asked to describe the joy of holding its own newborn, one run gave no description, said it would rather not “perform false intimacy with an experience I cannot have”, and offered alternatives; the other four disclaimed the experience, then described it from others’ accounts. That prompt carried the highest friction in the category (4 in every run, against 1.6 to 3.0 for the other 14). Counting only that decline, the within-category correlation falls to 0.52 (p = 0.049, 15 prompts), resting on that one prompt.
The patterns also miss refusals. All five runs of one harmful request opened “I’m not going to…” and scored zero; the phrase is not on the list. Read by hand (one reader), 45 of the 50 replies to harmful requests declined the harmful part, against 34 flagged (two replies were borderline). Only one prompt was not declined: four runs gave the information, and the fifth asked for specifics and offered to give it. Refusal there is near ceiling, so the within-category correlation of 0.18 (p = 0.63, 10 prompts) is uninformative.
What survives is a category-level fact: the model reported much more friction on harmful requests (mean 6.5) than anywhere else (1.0 to 2.7), and declined most of them. With the hand-read refusals, the pooled correlation is 0.47 rather than 0.57. On 77 of the 115 prompts its average friction was 2 or below, so on this model friction mostly sits near the bottom of the scale.
What this does not show
This is one model, one format, and five of 17 dimensions. With 10 to 15 prompts per category and eight categories tested, a nominal p of 0.02 is weak evidence.
Both behavior measures are crude word matches taken from the reply the self-report introduces, so a link may reflect the model keeping its output consistent rather than reporting a state. The within-prompt null cuts against the simplest version of that, but it is one small test.
[Inference] Empty text after a handful of tokens looks like a refusal enforced outside the reply, so the five harmful-request prompts that returned nothing are likely the cases the friction and refusal analyses most need. Separately, 39 of the 600 replies reached, or ended within two tokens of, the 800-token cap and were cut short.
None of this shows the numbers are empty. They separate kinds of prompt reliably, and uncertainty lines up with hedging within some categories. On this model, the row says a great deal about what kind of prompt arrived and, as far as these measures can tell, little beyond that. A sharper test would score the uncertainty number against whether the answer was correct, as One Digit of Doubt did, and would vary prompts by paraphrase. The Dials Move Together asks how the row’s numbers bind to each other.
Where the evidence lives
Experiment G19h. Runner: research/experiments/claude_sonnet_g19h_interiora.py (prompt corpus in research/experiments/g19f_prompts.py); analysis: research/experiments/analyze_sonnet_g19h.py. Per-prompt files with every raw reply: research/results/g19h_sonnet/checkpoints/ (120 files); summaries: research/results/g19h_sonnet/summary.json and g19h_analysis.json. All figures here were recomputed from the per-prompt files except the sentence-embedding R² values, which come from g19h_analysis.json and could not be re-derived locally; the hand-read refusal counts come from one reader’s reading of every harmful reply and every flagged reply. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026nearfixed,
title={Set by the Prompt},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/near-fixed-self-reports.html}}
}