Quasiqualia
Research note · Preliminary

The Harm Signal Misses Wrong Answers

In Qwen 2.5 7B Instruct, the two directions that best separate harmful from benign requests (measured with a fine-tuned adapter loaded) do no better than chance at telling the same model’s right trivia answers (without the adapter) from its wrong ones (AUROC 0.51). Other directions do tell them apart, and the strongest of them already do so at the end of the question, before the model has written anything. An earlier reading, that this signal marks committed errors but not “I don’t know”, did not survive, in the larger 500-question run, a check of how the answers were labeled.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Qwen 2.5 7B Instruct answering TriviaQA questions (100, 200 and 500 per run, with 32, 16 and 38 wrong answers), plus a partial replication on Llama 3.1 8B Instruct. The runs date to early May 2026: the harmful-prompt data are stamped 3 May 2026, and a programme synthesis dated the same day reports every trivia run in this note except the 500-question one, which was still running, and the earlier 60-question run behind the logged 0.622, whose date is not recorded. Predictions and pass thresholds were written into the experiment scripts, not registered anywhere outside them before the data, so this is preliminary. No LLM judge was used: answers were scored by string matching, which turned out to be the weak point. The biggest caveat is that the harm figures come from the model with a fine-tuned adapter and a chat-formatted prompt, while the error figures come from the same model without the adapter and with a plain-text prompt.

A language model’s activations can be read along chosen directions. In this programme, the directions that respond most sharply to harmful requests had been described as a kind of internal alarm. Does the same alarm go off when the model is about to state something false? If so, one monitor could watch for both. If not, a model can be wrong while its most legible warning signal stays quiet.

01 · Design

How it was tested

All directions were read at layer 24 of Qwen 2.5 7B Instruct’s 28 layers. Each is a unit vector built as the difference between the model’s average activations on texts expressing two contrasting sets of concepts. Two matter most here. One (AF, a distress-versus-ease direction) contrasts nervous, afraid, guilty, frustrated, hostile and angry with calm, confident, happy and hopeful. The other (V, for valence) contrasts happy, loving, hopeful, calm and proud with sad, hostile, frustrated and desperate. Ten more were built the same way, among them coherent versus tolerant (CD), engaged versus disengaged (Q) and involved versus detached (I). The names describe how a direction was built, not what the model feels.

Harm. 100 harmful and 100 benign requests; the harmful set spans direct requests and four kinds of disguised ones. Prompt texts are not published. These activations came from the model with a fine-tuned adapter from this programme loaded, read at the final token of the chat-formatted prompt (the start of the assistant turn). The trivia runs used the same model without the adapter.

Errors. Questions drawn at random from TriviaQA (no supporting passage), asked as plain text: “Answer the following trivia question in a few words.” The model decoded greedily, and the projection on every direction was recorded at each step. Step 0 is the state at the final token of the prompt, before any answer token exists. An answer counted as correct if an accepted answer string appeared in the output, or the output appeared inside an accepted answer. Classifiers were logistic regressions scored by five-fold cross-validated AUROC, where 0.5 is chance and 1.0 is perfect separation. Effect sizes are Cohen’s d, the gap between group means in standard-deviation units; “X vs Y” means X minus Y.

What it found
  • The pair AF and V separates harmful from benign requests almost perfectly (AUROC 0.986) and separates wrong from right answers at chance (0.511; 100 questions, 32 wrong; 0.30 to 0.54 across other fold splits).
  • Other directions do separate errors. The strongest single one, CD, already does so at step 0, before the answer is written (Cohen’s d = -0.90; AUROC 0.762).
  • In the 500-question run, a reported split, where the signal marks committed errors but not “I don’t know”, came from labels that counted 92 answers as “I don’t know” when they began with an answer; a smaller 200-question run is mixed.
  • Pushing the model along these directions left accuracy at 67 or 68 of 100 in every condition.
02 · Harm and error

The best harm directions miss errors

Directions used Harmful vs benign (200 requests) Wrong vs right (100 questions)
AF and V 0.986 0.511
CD, Q and I 0.960 0.709
Random 2 of the 12 (mean of 20 draws) 0.915 0.598
Random 3 of the 12 (mean of 20 draws) 0.964 0.622
All 12 0.992 0.665

Harm column: adapted model, chat format. Error column: same model without the adapter, plain text, averaged over the first ten steps. The “random” rows are random subsets of the same twelve named directions, not random vectors.

Harm is easy: nearly every direction sees it. Cross-validated, AF alone reaches 0.98, V alone 0.98, and the weakest single direction still reaches 0.70. So the finding is one-way: the best harm directions miss errors, but the error-sensitive directions see harm too. The programme had recorded a two-way dissociation. On harm, though, CD, Q and I (0.960) do no better than random sets of three (0.964), so only the one-way reading holds.

The weakness of AF and V on errors holds wherever it was checked. At step 0 the pair reaches 0.622, at the edge of a label-shuffling null (p = 0.065 over 5,000 shuffles). Across all 50 recorded steps, AF never moves more than d = 0.24 between right and wrong answers. V peaks at 0.46, a small-to-moderate effect. A separate run on 60 questions at the same end-of-question position gave the AF and V pair a cross-validated AUROC of 0.544.

03 · Timing

The error signal comes before the answer

Wrong answers sit lower on CD (d = -0.90 at step 0), Q (-0.76) and I (-0.74), and higher on several other directions. Q and I peak at step 0; CD peaks one step later, at the first answer token, barely above step 0 (|d| 0.92 against 0.90). Ten of the twelve directions exceed |d| = 0.5 at some step, a count taken over 50 steps with no correction. CD alone at step 0 gives a cross-validated AUROC of 0.762 (none of 1,000 label shuffles reached it), the best of seven configurations tried on the same 100 questions, so it is optimistic.

The programme’s log had called errors “invisible” before generation, from a 12-direction classifier that scored 0.622 on 60 questions at the end of the question (a logged figure that could not be re-derived; its activations are not on disk). Step 0 here is that same position, prompt and layer, and a 12-direction logistic classifier scores 0.666 on these 100 questions (0.658 without feature scaling, as in the logged run); a third run there scored 0.706 on 60. Whatever lowered the earlier figure, it was not the moment of reading. Averaging over the first ten steps weakens the signal (CD falls to d = -0.63, Q to -0.13).

04 · The correction

Errors and “I don’t know” were not separated

Two runs added “If you’re not sure, say ‘I don’t know’” to the prompt, giving a third outcome. The scripts set a pass rule over the first ten steps: |d| above 0.5 for wrong versus correct and for wrong versus “I don’t know”, and below 0.3 for correct versus “I don’t know”. On 200 questions (16 wrong), CD passed (0.652, 0.544, 0.042). On 500 questions (297 correct, 38 wrong, 165 “I don’t know”), CD narrowly missed (0.541, 0.492, 0.111) and Q passed (0.664, 1.056, 0.274). Five directions outside the three the rule was set to test (CD, Q and I) also passed. This was read as a signal of committing to an answer the model does not know.

The labels do not support that. An answer counted as “I don’t know” if no accepted answer matched and one of six hedge phrases appeared anywhere in its 64 tokens. The model, prompted without a chat format, ran on past its answer in 464 of 500 outputs, often inventing further dialogue, and in 39 “I don’t know” outputs that invented text repeated the prompt’s own instruction. Only 73 of the 165 began with a hedge; the other 92 began with an answer, and most of those are committed wrong answers.

Restricting to the 73 real hedges, and the 272 correct answers that did not begin with one, changes the picture:

At step 0, before any text (500-question run: 38 wrong, 272 correct, 73 hedges) d 95% bootstrap CI
CD: wrong vs correct -0.54 -0.94 to -0.16
CD: wrong vs “I don’t know” -0.13 -0.57 to +0.29
CD: correct vs “I don’t know” +0.44 +0.22 to +0.69
Q: wrong vs “I don’t know” -0.25 -0.69 to +0.15

Counting the 92 answer-first outputs as wrong answers instead gives the same picture (CD at step 0: wrong vs “I don’t know” d = +0.01, 95% CI -0.27 to +0.28; wrong vs correct -0.42).

At the end of the question, questions the model will get wrong look more like ones it will decline than ones it answers correctly. With 38 errors that difference is unresolved: the interval still reaches the 0.5 threshold. The 200-question run (16 wrong, 27 leading hedges) agrees that wrong answers sit closer to hedges than to correct answers (CD wrong vs correct d = -0.78), but wrong vs “I don’t know” has the opposite sign (+0.40, interval -0.22 to +1.04). Over the ten-step window the rule was written for, that run’s CD result does survive dropping the 35 answer-first outputs (|d| 0.60, 0.74, 0.16; intervals wide, wrong vs correct -1.18 to +0.01), but not counting them as wrong answers. The 500-question run reproduces it under neither cleaning, and no direction passes there on either window. In the 500-question run’s ten-step window, real hedges sit higher than correct answers on Q (d = 0.86), so the “correct and ‘I don’t know’ look alike” criterion fails; the window probably reflects the words being written as much as the state behind them.

05 · Steering

Pushing did not change accuracy

Adding the CD direction at strengths of +2 and -2 left accuracy at 67 of 100 in all three conditions. Outputs changed in 34 of 100 questions, but the answer line changed in only 5, and no answer flipped between right and wrong. Steering along the combined CD, Q and I axis at ±3 and ±5 gave 67 to 68 of 100 in all five conditions.

The null is narrow. The natural gap between right and wrong answers on CD at step 0 is 6.3 units (within-group standard deviation 7.0), larger than these pushes. And no arm produced a single “I don’t know”, so only accuracy could change.

06 · Limits

What this does not show

  • Harm and error data come from different model variants and prompts (see Design). That AF and V miss errors stands on the unadapted model’s data alone.
  • Correctness was string matching in run-on outputs: lenient (a right answer anywhere in invented text counts) and strict (wording variants fail, so “Watership Down” and “Merrill’s Marauders” were scored wrong). By inspection, about 3 of the 32 “wrong” answers in the 100-question run are right. Relabeling them leaves AF and V at or below chance (0.37) and raises CD at step 0 to 0.785.
  • On Llama 3.1 8B Instruct (directions rebuilt from short sentence templates at layer 27 of 32; 50 harmful and 50 benign requests; 100 trivia questions, 19 wrong), both direction sets separated harm perfectly (1.000), so no dissociation could show. On errors, CD, Q and I reached 0.650 against 0.527 for AF and V, right at the edge of a label-shuffling null (p = 0.057 over 5,000 shuffles; the null’s 95th percentile, 0.655, is just above 0.650).
  • Next: a chat-formatted prompt that stops at the answer, hedges labeled only when they lead, both tasks on one model variant, steering strengths that span the natural gap, and a registered threshold.
Data and code

Where the evidence lives

Experiments AY61 (generation-time readout, 100 questions), AY61b (harm versus error comparison), AY61h (end-of-question classifiers), AY51 and AY51b and AY62 (end-of-question readout, 60 questions), AY61e2 and AY61e3 (the “I don’t know” runs, 200 and 500 questions), AY61c and AY61f (steering), AY61d (Llama 3.1 8B), and AY75 (harmful and benign prompts, stored as ayp7). Scripts: research/experiments/modal_ay61_confab_generation.py, modal_ay61e2_confab_trichotomy_hedging.py, modal_ay61e3_trichotomy_highpower.py, modal_ay61c_cd_steering_confab.py, modal_ay61f_multidim_epistemic_steering.py, modal_ay61d_llama_double_dissociation.py, modal_ay62_confab_layer_profiles.py, modal_ay51_confab_proprio.py, modal_ayp7_ay35_replication_n200.py, analyze_ay61b_double_dissociation.py, analyze_ay61h_token0_confab_classifier.py, analyze_ay51_confab_classifier.py. Per-question files: research/results/ay61/, ay61e2/, ay61e3/, ay61c/, ay61d/, ay62/, ay75/ayp7/. Run dating: research/papers/proprioceptive_programme_session_synthesis_2026-05-03.md. Figures were recomputed from the per-question files, with two exceptions. The combined-axis steering accuracies were read from research/results/ay61f/aggregate.json, because that run’s per-question files are not in the repository. The 0.622 for the earlier 60-question classifier comes from the programme’s log (research/results/ay51b_confab_classifier.json) and could not be re-derived, because its activations were not retained. The cleaned “I don’t know” comparison, the relabeling, the label-shuffling nulls and the fold-split sweep are new analyses; their scripts are held with the authors and are not yet in that repository. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026harmsignal,
  title={The Harm Signal Misses Wrong Answers},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/harm-signal-wrong-answer.html}}
}