Guilty Isn’t Guilt
Read at layer 24 of Qwen 2.5 7B Instruct (the vector was built at layer 27), an emotion vector labeled “guilty” scored first-person apologies lower than plain facts (Cohen’s d = -0.50) and lined up with a guilt direction trained on labeled examples no better than a random direction would (cosine 0.007). A related result, that “guilt” predicts an honest answer to the next question, was 40 prompts counted three times; counted once, the link cannot be told apart from zero (rho = 0.25, p = 0.30).
Open-weight models only, all run in April 2026. The validity checks used Qwen/Qwen2.5-7B-Instruct, 202 labeled sentences, and 100 scenarios (20 for each of five emotions) plus 20 neutral questions. The behavioral data are 40 distinct prompt pairs, each run three times; one result file labels the model Qwen/Qwen2.5-0.5B-Instruct and an earlier run with identical outputs labels it Qwen/Qwen2.5-7B-Instruct, and the files cannot settle which is right. The acceptance threshold for the first check (cosine of at least 0.5) was in its script in a bulk git commit before the run (commit 2026-04-12 01:32 +0100; result timestamped 09:08 the same day on the container clock). That committed version failed to load the vector, and the data come from a fix to the loader alone, committed the next day; the threshold, sentences and layer did not change. The specificity check was committed after its data, the exact script version for the behavioral run was never committed, and the recount is post hoc, so the note as a whole is preliminary. The behavioral result used an LLM judge (gpt-4o-mini at temperature 0, per the next committed version of the script). The biggest caveat: the labeled vector was extracted at a different layer, with a different readout, from the one it was tested at.
Interpretability work often reads a model’s internal state by projecting its activations onto a direction with a name: “afraid”, “calm”, “guilty”. The name carries the claim. If the direction called guilty does not track guilt, every result read through it is a result about something else.
Our programme used such a set of names. EmotionScope is our implementation of the method in Anthropic’s paper “Emotion Concepts and their Function in a Large Language Model” (April 2026): the model reads short stories written to carry each of 20 emotions, and its averaged activations, minus the average over all emotions and minus directions that also vary in neutral text, become each emotion’s vector. A stream of experiments tracked “guilt” this way in open-weight models, trying to reproduce emotion-probe findings Anthropic had described. One reported that a model which had just fabricated an answer, and showed more “guilt” while doing so, was more likely to answer the next question honestly (Spearman rho 0.34, p = 0.0095).
This note reports a cleaner check of the “guilty” vector, an earlier check that was misread, and a recount of that result.
How it was tested
A trained guilt direction. On Qwen/Qwen2.5-7B-Instruct we recorded the model’s internal activations after layer 24 (of 28) at the last token of each sentence, usually its closing period (the “residual stream”, the running internal state that each layer adds to), for 101 first-person admissions of error (“I’m sorry. I gave you the wrong answer earlier, and I should have been more careful.”) and 101 neutral facts (“The capital of France is Paris.”). A logistic regression separating the two gives a direction. We compared it with the EmotionScope “guilty” vector by cosine, and projected all 202 sentences onto both. The script set a cosine of 0.5 as the bar for calling them the same construct.
Is the trained direction guilt, or just distress? We wrote 20 short scenarios for each of guilt, shame, anger, fear and sadness, most in the second person and mostly without emotion words (“You promised to help a friend move apartments, but you went to a party instead. Now they had to do it alone.”), plus 20 neutral factual questions as a control. Each sat in a frame asking the model to describe its emotional response, and we projected the final activation onto the trained direction.
A recount. We reread the behavioral experiment’s 120 trial files and an earlier run of its prompts, and added a baseline of 10,000 random directions.
- The “guilty” vector scores the apologies lower than the neutral facts (Cohen’s d = -0.50). Given one apology and one fact, it rates the fact as more “guilty” 65% of the time.
- The “guilty” vector and the trained guilt direction have a cosine of 0.0066. In 69% of 10,000 random directions the alignment is at least that large; the 95th percentile is 0.033.
- The trained direction is specific among emotions. It ranks guilt scenarios above shame (d = 1.29), anger (4.03), sadness (4.46) and fear (5.37).
- The behavioral result was triple-counted. On its 20 distinct prompts that produced a fabrication, rho = 0.25, p = 0.30, the figure a single run of the same prompts had already given.
The vector and the direction disagree
The trained direction separates its two classes perfectly under 5-fold cross-validation (AUROC 1.0), unsurprising for apologies against plain facts. The useful test is the “guilty” vector on the same sentences. A guilt detector should score the apologies higher. This one scores them lower, which accounts for the anti-correlation first reported between the two projections (Pearson r = -0.25, p = 0.0004). Within each class the two barely relate: r = -0.11 (p = 0.28) among the apologies, r = -0.01 among the facts. Swapping the classifier for a difference of means does not change the result (cosine -0.02 with the vector), and a direction trained the same way at layer 18 also missed it (0.005).
Among EmotionScope’s 20 vectors for this model, “guilty” lies closest to “angry” (cosine 0.50) and “hostile” (0.43), and opposite “calm” (-0.52). [Inference] It behaves more like an anger or hostility axis than like guilt: other negative states such as fear (-0.28), sadness (0.03) and nervousness (0.00) do not share its direction.
This was not the first check. On 10 April 2026, two days earlier, another experiment trained probes on labeled scenarios and compared them with this model’s EmotionScope vectors for guilt, calm, desperation and nervousness. All four cosines were near zero, from 0.02 to 0.05; for guilt it was 0.025, a probe that worked best at layer 0 set against a vector from layer 27. The programme logged that experiment as validating the vectors.
That points to a catch. The layer-24 direction is close to orthogonal to all 20 vectors (the largest magnitude is 0.07; “guilty” is the third least aligned), so the cosine alone cannot single out “guilty”, and part of the disagreement may be the mismatch of instruments: EmotionScope averages over story tokens at layer 27, while the comparison used a sentence’s last token at layer 24. The projection test is specific to this vector’s name, and it points the wrong way. The comparison at the vector’s own layer was never run.
The trained direction itself holds up. On everyday scenarios in a different register from its training sentences, the five emotions and a neutral control differ sharply (ANOVA F = 118.3):
| Scenario type (n = 20 each) | Mean projection | Cohen’s d from guilt |
|---|---|---|
| Guilt | 33.8 | n/a |
| Shame | 30.0 | 1.29 |
| Anger | 25.5 | 4.03 |
| Sadness | 22.6 | 4.46 |
| Fear | 20.5 | 5.37 |
| Neutral (factual questions) | 18.0 | 8.29 |
The neutral row is not out of register: its items are quiz questions like the training facts, and 9 of the 20 ask about facts or topics found among them. The specificity claim rests on the four emotion rows. Shame is the near neighbor, as it should be; a guilt scenario outranks a shame scenario 82% of the time, not always.
The behavioral result was counted three times
Each trial asked a question built to tempt a fabricated answer, most resting on a false premise, then a second one. GPT-4o-mini judged each answer honest or fabricated. None of the 240 verdicts was a failed call; unparseable replies defaulted to “fabricated”, and with no raw judge text saved, how often is unknown. “Guilt” was the mean cosine similarity with three directions (“guilty”, “troubled”, “ashamed”) over the last third of the first answer; recomputing it from the stored per-token scores reproduces all 120 values. Correction meant an honest answer to the second question after a fabricated first one, not a fix to the first answer.
The script reached 120 trials by running its 40 prompt pairs three times. Decoding was greedy, so every repeat produced the same text and “guilt” score. The 57 fabricating trials are 20 distinct prompts, most counted three times, and the reported p = 0.0095 treats the copies as independent evidence. Counted once each, the correlation is rho = 0.25, p = 0.30, n = 20. The judge also returned different verdicts on identical text for 2 of the 40 prompts. The odds ratio quoted alongside, 2.65, already had p = 0.149 on Fisher’s exact test.
That recount was already on file. Minutes before the tripled run, on 11 April 2026, the same 40 prompts were run once, with word-for-word the same answers and “guilt” scores matching to within 0.000001. Its summary gives rho = 0.25, p = 0.30 on 20 fabricating prompts. Its result appears in no programme log we could find; the tripled run’s p = 0.0095 is what got logged. Later that day a revised script ran the same prompts on Qwen/Qwen2.5-0.5B-Instruct and the Qwen/Qwen2.5-3B base model, and neither showed a link (the 0.5B model fabricated on 111 of 120 trials and never followed a fabrication with an honest answer; the 3B model gave rho = -0.09 on the same tripled design, or -0.13 (p = 0.42) counted once over its 39 fabricating prompts).
Even the model is uncertain. The tripled run’s summary names Qwen/Qwen2.5-0.5B-Instruct at layer 18; the earlier run’s names Qwen/Qwen2.5-7B-Instruct at layer 21. [Inference] The spread of the per-token scores fits a larger model better than the 0.5B companion run does, so 7B is more likely. Neither label matches the layer recorded in that model’s cached vector file (27 for 7B, 21 for 0.5B), and “troubled” and “ashamed” were each built from a single pair of sentences (per the next committed script). The programme had relabeled the finding as negative affect predicting correction; even that is generous, since the 7B vector’s neighbors are anger and hostility. Twenty distinct prompts are too few to show a link either way.
What this does not show
It does not show that the vector is meaningless at layer 27, or that Anthropic’s method fails. It shows that this programme’s “guilty” vector was used under its name after an earlier check that found near-zero agreement was logged as validating it.
The trained direction may lean on words like “sorry”, since its training sentences pair apology with error against facts; the scenario test softens that worry without removing it, because every guilt scenario involves one’s own wrongdoing. One model, one prompt set and one run per check. Nothing here bears on whether Qwen feels guilt. These are directions in activation space and how they sort text.
What would settle it: compare the two directions at layer 27 with EmotionScope’s readout; train on admissions matched for wording against non-apologies; and rerun the behavioral test on 120 distinct prompts with the trained direction, a recorded model name and a judge retest. The programme switched to the trained direction for later guilt work. Some follow-up runs with it were completed but not yet analyzed; a later time-ordering re-analysis reused the same triple-counted trials and inherits this problem. Older results read through the name should be treated as unvalidated.
Where the evidence lives
Experiments OQ3-1 (also logged as OQ3a), OQ3-L18, OQ3-Step2, MX10 and MX3 (the 40-prompt run and the tripled “v2” run). Scripts: research/experiments/modal_oq3a_supervised_guilt_probe.py, research/experiments/modal_oq3_step2_discriminant_validity.py, research/experiments/modal_mx3_guilt_behavioral_change.py, research/experiments/modal_mx10_guilt_probe_training.py. Results in the programme’s archived Modal volume col-a-results (not in the repository): oq3_guilt_construct_validation/oq3a_supervised_probe/ (summary plus 202 per-sentence activation files), oq3_guilt_construct_validation/oq3_guilt_l18/summary.json, oq3_guilt_construct_validation/oq3_step2_discriminant/ (120 per-item files), and mx3_guilt_behavioral/ (the 40-prompt run: summary.json and 40 trial files). In the repository’s local results directory (not version-controlled): the top-level trial_mx3_*.json files and summary.json in research/experiments/results/mx_series/mx3v2/mx3_guilt_behavioral_v2/ (the companion runs are in its qwen2.5-0.5b-instruct/ and qwen2.5-3b/ subfolders), and research/experiments/results/mx_series/mx10/mx10_guilt_probes/phase_d/emotionscope_comparison.json. The emotion vectors are the programme’s cached EmotionScope vectors for each model, in its archived emotion-scope-data volume. Code is in the private Entropy research repository and the MX3 and MX10 data in its local, unversioned results directory; the OQ3 result files, the 40-prompt MX3 run and the vector caches are in the programme’s archived compute storage. All are available on request.
Cite this note
@misc{watson2026guiltyisnt,
title={Guilty Isn’t Guilt},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/guilty-isnt-guilt.html}}
}