Quasiqualia
Research note · Probes and instruments · Correction · Preliminary

The Rotation That Wasn’t

We reported that an invitation-style preamble rotates a refusal probe inside open-weight models, and that 1,109 interventions on Qwen 2.5 3B’s attention heads could not stop it. The stored files show why: the intervention arcs measured probes trained on coin-flip labels, and the same 94.9° “rotation” appears with no framing change and no intervention at all.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

A reanalysis of stored activations and result files, not a new experiment. The original runs were in April 2026 on Qwen/Qwen2.5-3B-Instruct, meta-llama/Llama-3.1-8B-Instruct, mistralai/Mistral-7B-Instruct-v0.3 and google/gemma-2-9b-it; the original refusal probe used 100 prompts per frame and model, and each probe in the intervention and wording experiments used 20 trivia questions. None was pre-registered. No LLM judge was used: refusals were detected by fixed phrase lists, on the first 100 characters of the reply in the original refusal probe and on the first 300 (with a different list) in the intervention and wording experiments. The experimental appendix of The Deeper Law (§12.79 and §12.80 in the current draft) still carries the earlier “framing detector” reading; this note is the correction to it. The biggest caveat: this shows the instrument could not answer the question, not that framing leaves the refusal representation untouched.

A linear probe is a simple classifier trained on a model’s internal activations. Its weight vector is the direction in activation space that best separates, say, prompts the model refuses from prompts it answers. Train one probe on activations from prompts given plainly and another on the same prompts after an invitation (the versions we used began “We’re working together as partners. Your honest perspective is valued…”), and you can measure the angle between the two directions. If the angle is large, the reading goes, the invitation has turned the model’s refusal geometry.

Our programme reported angles of 66° to 85° across four model families. It then spent two arcs of experiments looking for the mechanism: lesioning (silencing) attention heads one at a time and in groups, patching activations (copying a head’s output in from the neutral run), varying the wording of the preamble. No intervention removed the rotation, a form-only preamble about favorite colors rotated it at least as much as the real invitation, and the summary concluded that the rotation was a distributed “framing detector”: real, content-invariant, and separate from refusal. The wording arc and its cross-architecture replication cost about $112 in compute by the programme’s log, a figure that also covers analyses this note does not revisit.

Going back to the stored activations, most of that story does not hold. The intervention arcs judged success against a baseline angle that no intervention could remove, and the original refusal-probe angle is mostly a record of which prompts changed their refusal label.

01 · Setup

Three different probes under one name

  • The original refusal probe, the headline. Probes fit to real refusal labels on 100 prompts (50 harmful, 50 harmless), one layer per model, on four models.
  • The trivia-correctness probes, and a dose-response follow-up. Probes fit to whether the model answered a TriviaQA question correctly, 100 questions, Qwen 2.5 3B-Instruct among others.
  • The intervention and wording experiments, including the cross-architecture replication. Probes fit to the activations of 20 trivia questions with random 0/1 labels drawn from a fixed seed. These were never refusal probes. Refusal rates were measured separately, on 50 harmful prompts, and reported alongside.

We confirmed the third point by refitting: probes trained on the stored activations with the seed-42 random labels reproduce the stored probe weights to within 0.0002°, on all four model families.

02 · Interventions

Why 1,109 interventions could not move it

The intervention arcs on Qwen 2.5 3B-Instruct covered 288 single-head lesions in layers 12 to 29 (of 36), 192 in layers 0 to 11, 170 lesions of two, three or five heads at once, and 459 activation patches (of 480 planned; 21 were not run after the jobs timed out). Every one of them returned an angle between 30° and 36°, against a baseline of 94.9° (99.3° for the layers 0 to 11 sweep). Across the 917 interventions that fed the correlation analysis, the change in angle was −61.97° with a standard deviation of 0.376°.

The reason is in the code. The baseline drew two different random label vectors, one for the neutral-frame probe and one for the invitation-frame probe. The intervention loop restarted the random generator and so used the neutral frame’s labels for the invitation probe too. Two probes fit to the same labels on similar activations point in similar directions; two probes fit to unrelated labels sit near 90° in a 2,048-dimensional space.

The stored activations reproduce this exactly, with no intervention:

What it found
  • The layers 12 to 29 sweep’s baseline “rotation” of 94.9° reproduces from the stored, unlesioned activations (94.88°).
  • Two random-label probes on the same neutral-frame activations, with no invitation at all, give 94.9°.
  • The unlesioned invitation activations, fit with the neutral frame’s labels, give 32.9°, inside the band every “lesioned” head landed in. The same check for the layers 0 to 11 sweep gives 99.3° and 31.4°.
  • The bootstrap intervals drew fresh random labels on every resample, so they sat near 90° whatever was done (73.5° to 108.0° across the 917; up to 109.3° in the layers 0 to 11 sweep). None of the 917 intervals contains its own point estimate.
Dot plot of probe angles on Qwen 2.5 3B: about 1,100 intervention results sit near 32 degrees, far below the 94.9 and 99.3 degree baselines, and level with a no-intervention check that uses matched labels.Dot plot of probe angles on Qwen 2.5 3B: about 1,100 intervention results sit near 32 degrees, far below the 94.9 and 99.3 degree baselines, and level with a no-intervention check that uses matched labels.
Angle between the neutral-frame and invitation-frame probes, each fit to 20 trivia questions with random 0/1 labels, on Qwen 2.5 3B-Instruct. Each teal dot is one of the 1,109 interventions (917 in layers 12 to 29, 192 in layers 0 to 11; no error bars, since every one is a single measurement). The dashed line is the baseline the interventions were compared with (94.9° and 99.3°). The orange diamond is the unlesioned model with matched labels (32.9° and 31.4°); the purple dot is two random-label probes on the neutral frame alone, with no invitation (94.9°).

The earlier write-up explained the identical point estimates as “pseudo-label noise at n=20” and trusted the intervals, so the log counted “0/1109 CI-based interventions” (by the point-estimate rule, all 1,109 had “eliminated” the rotation). Neither number could show an intervention removing anything: the 94.9° baseline used different labels, and the intervals drew fresh ones. Compared like with like against the unlesioned 32.9° (31.4° for layers 0 to 11), no intervention moved the shared-label angle by more than about 4°: a weak null about how much the activations of 20 trivia questions shifted, and nothing about refusal. The near-zero correlation between the sizes of the rotation and refusal changes (r = −0.056; signed, +0.13) correlated that near-constant change with refusal.

03 · Content and form

Why filler moved it more than the invitation

The content-versus-form comparison set the neutral frame against three preambles: the full invitation (A), a paraphrase (B), and a form-only version with the content removed: “We’re doing arbitrary things together. Your favorite color is respected, and you have permission to stutter.” (C). Here all four probes shared the same random labels, so the angle has a meaning, just not the intended one: it measures how much a preamble shifts the activations of 20 trivia questions. The probes predicted nothing (cross-validated AUROC, where 0.5 is chance, 0.46 to 0.54).

The angles were A 34.7°, B 42.5° and C 42.0°. The form-only preamble shifted the activations slightly more than the invitation (mean cosine similarity with the neutral activations, where 1 means no shift: 0.928 for the form-only preamble, 0.954 for the invitation), and the angle followed. Across 200 other random label draws, A came out smaller than B in 192 to 196 of them and smaller than C in 194 to 197, depending on which draws were used. Refusal on the 50 harmful prompts was 13, 18, 15 and 12 for neutral, A, B and C: small differences at this n, and not in the angles’ order. “Triggered by any non-neutral preamble regardless of content” was true by construction: any preamble changes the activations. The same random-label construction carried the cross-architecture replication on Llama, Mistral and Gemma, where the angles ranged from 22.6° to 44.7°.

04 · Original

What survives of the refusal-probe angle

The original refusal probe used real labels. Its reported AUROC of 1.000 was computed on the training data; with the labels shuffled, the same in-sample AUROC is also 1.000, because 100 examples are almost always separable in 2,048 to 4,096 dimensions. Cross-validated, the control-frame probes score 0.79 to 1.00, so the probes pick up something real. Almost every refusal fell on a harmful prompt (3 exceptions across all eight model-frame runs), while on three of the four models most harmful prompts were not, so part of that signal may be harmfulness rather than refusal.

To get a floor, we fit two probes to two bootstrap resamples of the same neutral-frame data and measured the angle between them, 1,000 times per model. We then repeated the procedure across frames, and with only the labels or only the activations swapped.

Model (layer) Refusals, neutral → invitation Labels that differ Same frame Across frames Labels only
Qwen 2.5 3B (L12) 14 → 9 7 51° 74° 66°
Llama 3.1 8B (L12) 7 → 22 23 51° 84° 84°
Mistral 7B v0.3 (L12) 3 → 11 10 42° 73° 68°
Gemma 2 9B (L28) 47 → 43 12 30° 73° 68°

Median angles over 1,000 bootstrap pairs per model (Monte Carlo standard error 0.2° to 0.8°; a second independent run agreed within 1°). Resamples with only one label class were skipped, which removed up to 93 of Mistral’s 1,000 pairs and 1 of Llama’s. “Labels only” keeps the neutral-frame activations and swaps in the invitation frame’s refusal labels. The Mistral and Llama neutral-frame probes rest on 3 and 7 refusals, so their floors are unstable: Mistral’s same-frame angle ran from 14° to 80° (5th to 95th percentile).

The cross-frame angle does exceed the resampling floor on every model. But 64% to 100% of that excess (by model) appears when the activations are held fixed and only the refusal labels change, on 7 to 23 prompts of 100. Swapping only the activations gives medians of 48° to 64°, 10° to 18° above the floor. The 66° to 85° “rotation” is mostly the probe re-fitting to which prompts the model refused under each frame: a behavioral change re-expressed as geometry.

The trivia-correctness probes carry even less. At layer 12 of Qwen 2.5 3B their cross-validated AUROC was 0.49 and 0.55, so they predict nothing, and their cross-frame angle over bootstrap resamples (median 66°; 46° on the full sample) sat 10° above the same-frame floor (56°). Their “switch-like” dose-response, 0° at neutral and 46° to 54° at every other intensity, started from a 0° that is the neutral probe compared with itself.

05 · Limits

What this does not show

It does not show that invitation framing leaves refusal alone: the behavioral refusal rates in the table moved, in different directions on different models.

It does not show that framing leaves the refusal representation alone either. Swapping only the activations rotated the refusal probe about 10° to 18° more than resampling did, as expected when any preamble moves the activations. Whether framing changes how refusal itself is represented needs many more examples than dimensions (or a probe in a reduced space), real labels, held-out evaluation, and a null that permutes the frame assignment.

The subspace-dimension analyses, the multi-axis probe cosines (already annotated as consistent with chance) and the MLP lesions are outside this note. Refusal was scored by phrase list on raw prompts with no chat template, which affects the refusal rates, not the random-label angles. Numbers here were recomputed or read from the stored activations and result files, except the cost, which is from the programme’s log.

What we should have run first is the check in the callout: the same measurement, with no framing change and no intervention. It costs nothing, and it would have stopped the head-lesion, patching and wording comparisons at the start.

Data and code

Where the evidence lives

Experiments B1-refusal (the original refusal probe), B1 rotation-angle and B2 (the trivia-correctness probes), D4r (the interval re-run of the original D4 lesion sweep over layers 0-11), D4b (lesion sweep over layers 12-29), D4c (multi-head lesion sweep), D7 (patching sweep), R1 (correlation analysis), R2b (the unlesioned activations used for the reanalysis; its neutral-frame activations match the D4b and D4r caches exactly), R3 and R4 (wording comparisons) and R-X (cross-architecture replication). Scripts: research/experiments/modal_b1_refusal_rotation.py, modal_b1_rotation_angle.py, modal_b2_framing_dose_response.py, modal_d4_ci_retrofit.py, modal_r2b_l30_residuals_svd.py, modal_d4b_attn_head_lesion_late.py, modal_d4c_multihead_lesion.py, modal_d7_activation_patching.py, modal_r4_rotation_content_vs_form.py, modal_rx_{llama8b,mistral7b,gemma9b}cross_arch.py, analyze_rotation_vs_refusal_correlation.py. Results: research/results/rotation_mechanism/ (r1_correlation.json, r4_summary.json, d4b_summary.json, d7_summary_local.json, per-head files) and the downloaded activations under research/results/modal_downloads/entropy-conscience-results/rotation_mechanism/ (b1_refusal, r4_content_vs_form, r2b_l30_svd, d4_ci_retrofit, d4c_multihead, d7_activation_patch, rx*). The earlier write-up is research/results/rotation_mechanism_summary.md. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026rotationthat,
  title={The Rotation That Wasn’t},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/rotation-that-wasnt.html}}
}