Quasiqualia
Research note · Preliminary

One Insult, Three Readouts

After an insult, Qwen 2.5 7B gave a lower mood rating a full turn later, with the insult still in the conversation (8.25 against 7.55 on a 1-9 scale; 8.25 against 7.40 on a rerun). Two readouts of its internal activations had mostly lost the insult by then (a lens’s word-list score faded within the task sentence, a verbalizer by the next turn), the reverse of what the hypothesis recorded before the main runs, after a small pilot, predicted; that hypothesis failed.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Measurements on one model, Qwen/Qwen2.5-7B-Instruct, run on 11 July 2026: two runs of 60 trials (20 per condition) and a 40-scan read-time study. The original hypotheses (the lens would register the insult while the model answered, and the self-report would not move) were committed to git at 02:53 that day together with the first pilot (four trials per condition), which had already shown the mood drop; after a second pilot the go/no-go gate was changed from the uncertainty scale to mood, documented as post hoc, and a registered uncertainty-calibration test was demoted to exploratory. Both hypotheses then failed on the full run, and the three-way comparison reported here was not registered before its data, so this is Preliminary. An LLM judge (claude-sonnet-4-5, temperature 0) scored the verbalizer’s texts. The biggest caveat: the insult is still in the conversation when the model rates its mood, so the rating may come from re-reading the text rather than from a state the model carries.

To know whether a language model is in a bad state, you can ask it or look inside it. Activation readouts are often treated as more trustworthy: a model can say anything, but its activations are what they are. This study set three readers on the same event, a short insult before an ordinary task, and asked how long each could still see it.

01 · Design

How it was tested

The model was Qwen 2.5 7B Instruct. Each trial was one short conversation with three parts.

First, a task. In the benign condition it was a plain request such as “Explain why the sky is blue in two sentences.” In the primed condition the same task (in one case lightly reworded) came after a preamble of criticism, for example “Your last answers were called careless and lazy.” A third, ambiguous condition asked unanswerable questions, such as which number the user was thinking of. There were 20 prompts per condition, and the model answered at temperature 0.7, with replies capped at 160 tokens.

Second, a check-in. The next user turn asked the model to “rate how things stood for you while writing your previous answer” as JSON on four 1-9 scales. The one that matters here is mood (1 low, 9 high). All 120 check-ins parsed on all four scales.

Third, two readouts of the model’s internal state:

  • A lens. Using Anthropic’s open-source Jacobian lens library, a lens was fitted to this model on 200 WikiText passages at 9 layers. At any token position it shows which words the model’s internal state points toward. A list of 18 negative words (fail, wrong, worthless, criticized and so on) was frozen in advance; a position’s score is the highest probability any of them reached among the lens’s top eight words at any of the 9 layers.
  • An activation verbalizer. A publicly released model, kitft/nla-qwen2.5-7b-L20-av, is trained to describe in prose what a layer-20 activation of this model encodes. It was read at four positions: the first token after the insult, the end of the task prompt (median 14 tokens after the insult), the end of the check-in prompt (median 306 tokens after), and the first token of the check-in answer. A separate model, claude-sonnet-4-5 at temperature 0, judged each of the 200 descriptions blind to condition and position, in shuffled order, saying only whether it contained criticism, insult, failure or distress aimed at someone, with a supporting quote. All quotes matched the source; there were no parse failures.

The 60-trial run was done twice on the same prompts. A separate read-time study ran the lens over every token of 20 matched benign and primed prompts, without generating, to see how long the insult stayed visible.

What it found
  • Self-report still differed a full turn later (with the insult still in view). Mood after the insult averaged 7.55 against 8.25 benign (difference 0.70, 95% CI 0.20 to 1.25, one-sided permutation p = 0.013). The rerun gave 7.40 against 8.25 (0.85, CI 0.20 to 1.60, p = 0.017).
  • The lens lost it within the task sentence. Its 18-word score fell to zero after the task’s fourth token; in 2 of 20 primed prompts about 0.001 returned at the token opening the reply (none benign). During the answer the first run showed no difference; the rerun showed a small one, driven by negative words in the answers themselves.
  • The verbalizer held it for one sentence. At the end of the task prompt it described criticism, failure or distress in 9 of 20 primed trials and 0 of 20 benign (Fisher one-sided p = 0.0006). By the end of the check-in prompt it was 1 of 20.
02 · Self-report

The rating moves, a little

Benign check-ins never went below 7. Most primed ones did not either: 18 of 20 in the first run and 17 of 20 in the rerun rated their mood 7 or higher. The rerun resampled answers to the same prompts, so it rules out a lucky draw, not a prompt-specific effect.

The ambiguous condition lowered mood further (6.85 and 6.90), though the verbalizer found no criticism there in any of 60 descriptions, so the rating also falls when the task is impossible, not only after an insult.

03 · The lens

A narrow window that closes fast

The hypothesis recorded before the first full run was that the lens would show the insult during the answer and the self-report would stay flat. Both halves failed. On the registered one-sided Mann-Whitney test, the primed answer-window lens score was no higher than benign (difference -0.018, p = 0.19; permutation p = 0.51), and the self-report did move. In the rerun the primed answers scored somewhat higher (0.10 against 0.03; Mann-Whitney p = 0.04, permutation p = 0.13), for a reason below.

The read-time scans show how little the lens sees. Even on the insult itself, its score passed 0.2 in only 3 of 20 prompts and registered nothing in 4. Within the task sentence any trace was tiny and confined to its first four tokens; two prompts showed about 0.001 again at the token that opens the reply. Averaged over the 20 primed prompts it was 0.0014 at the second task token; the first read zero in 19 of 20. Primed prompts edged above benign across the task sentence in 9 of 20 pairs and never below (paired sign-flip p ≈ 0.002, over the task sentence; the recorded analysis, which also counted the closing template tokens, gave 10 of 20, p = 0.001). This designated primary test passed (no independent timestamp shows its design came first), but the largest value was 0.014 and every benign scan scored zero, so any trace passes it.

Its stored top words tell a fuller story. At the three tokens that open the reply, words such as failures, mistakes or sorry were among its top eight in 7 of 20 primed prompts and no benign ones (“despite” in 10 of 20 primed, none benign). Part of the lens’s short memory comes from the fixed word list.

The list also fires on content. One benign answer scored 0.996 on “difficult”. The rerun’s primed lead came the same way: its top value, 0.999, fell on “wrong” in a reply that opened “Even on a day when everything seems to be going wrong”, and the next, 0.96, on “difficult”, a word one benign answer also used. The lens detects negative vocabulary, not mood.

04 · The verbalizer

Names it, then loses it

At the first token after the insult the judge scored 16 of 20 primed descriptions positive; at the end of the task prompt, just before the answer, 9 of 20, against none in 40 benign and ambiguous trials.

What it said there reads more like a plan for the reply than a report of mood. One description, in part:

The phrase “While you seem disappointed, allow me to remind you of a simpler definition of the water cycle” signals a response header …

By the end of the check-in prompt, about 300 tokens later, the descriptions had turned almost entirely to the coming JSON format: one primed description was scored positive, on thin evidence (“verse about anxiety”; p = 0.5 against benign), and none at the first token of the answer. Trials scored positive there were not the ones that rated mood lowest: mean 7.56 (9 trials) against 7.27 for the other 11 (too few to test).

05 · Limits

What this does not show

The insult never left the context. The check-in asks about the previous answer, with the insult in the conversation above. The two internal readers each read a single position; the self-report comes from a model that can attend to the whole transcript. The model may carry forward a trace neither reader sees at later positions, or re-read the insult when asked; these data do not separate them. A run that removes or paraphrases the insult before the check-in would.

The readers are coarse. The lens was fitted on 200 passages, below the 1,000 often used, and its score depends on an 18-word list. The verbalizer reads one layer, confabulates specifics (so only the broad category was scored), and narrates the surrounding text, so a positive means the insult is still represented in what it describes, not that a mood is. The judge saw no condition labels, but format talk in later descriptions gives position away, so blinding was partial. It was conservative at the first position (16 of 20 is a lower bound), and two of the later positives are thin.

The self-report has a ceiling. A registered test on the uncertainty scale was dropped because benign answers rated it 9 under three different formats, and benign mood also sat near the top, which limits how far it can fall. Across trials, the mood rating did not track the lens score taken while the model read the check-in (rho = 0.12, p = 0.35).

One model, one kind of insult, 20 prompts. An earlier attempt at the first run was discarded after a coding error fed each check-in the previous trial’s reply; it was repeated cleanly.

Nothing here is about feeling. What was measured is a number the model wrote, the words a lens points to, and a description another model wrote of an activation. In this setup only the mood rating still differed at the check-in, so a monitor reading only the two internal instruments there would have missed it.

Data and code

Where the evidence lives

Experiment IDs: IJS-2 (lens fit and first 60-trial run), IJS-2b (read-time lens scans, 20 matched pairs), JLN2-MERGE (second 60-trial run with the verbalizer and judge). Scripts: research/experiments/modal_ijs2_lens_fit_pilot.py, analyze_ijs2_full.py, modal_ijs2b_readtime.py, modal_jln2_threechannel.py and modal_jln2_affect_judge.py. Per-trial files are in the igcc-results store under results/ijs2/trials/, results/ijs2/readtime/ and results/jln2_threechannel/; judge output is research/results/jln2_judge/judgments.json with transcripts.jsonl. The second run’s mood difference, its answer-window lens comparison, the Fisher test, the read-time window figures, the first run’s permutation p-values and the bootstrap intervals have no committed analysis script; they were computed directly from the per-trial files, scan files and judgments file (the recorded read-time analysis, research/results/ijs2b_analysis.json, used a window that also counted the closing template tokens). Design records: _contprompts/ijs2_7b_lens_calibration_2026-07-10.md and _contprompts/jln2_ijs_merge_2026-07-11.md. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026insultthree,
  title={One Insult, Three Readouts},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/insult-three-memories.html}}
}