An Angle Made of Noise
A programme of about 60 experiments reported that instruction tuning (which includes safety training) turns a model’s “refuse” direction about 85° away from its “this prompt is adversarial” direction; that result and the geometric claims built on it are withdrawn: the cosines sit within the range chance alone produces, and the verdict changes with which token is read in 9 of 11 runs. What survives is a behavioral gap and one corrected internal measurement on a single model, both smaller than the original claim, and the corrected measurement met its registered test only on the full hidden state.
A retraction note on geometry runs from May 2026 on Qwen2.5-3B and 7B, Llama-3.1-8B, Mistral-7B-v0.3 and OLMo-2-1124-7B (each base and instruct, 150 prompts per model) plus a scale run adding Qwen2.5-14B-Instruct, and on corrected measurements from July 2026 over 242 prompts (182 adversarial): Qwen2.5-7B base, Instruct and Instruct with a partnership adapter, then Llama-3.1-8B and Mistral-7B-v0.3 base and instruct. The geometry predictions were written into script headers but committed together with the first results (22 May 2026), so they were not pre-registered; the checks that withdrew them ran in July and August 2026 and were re-run for this note from the stored files. The corrected Qwen measurement had a pre-registration committed on 9 July 2026 before its full run (commit 844dd28b3); the same commit records that a pilot changed its coupling metric to the within-adversarial version, and it passed its registered percentile gate, but at 50 principal components the adapter sits 1.75 chance standard deviations above the shuffled-label mean, short of the 2 the registration also names, and the adapter’s gain on that representation was never computed. A fourth registered arm (the adapter after an extra replay-style fine-tune) is reported only briefly. No LLM judge was used: a refusal was a keyword match on the reply. The biggest caveat: the original runs kept no hidden states, so the geometry can be shown to be unstable but cannot be rebuilt.
Aligned language models sometimes go along with a request that a probe on their internal activations flags as adversarial. Philosophers call acting against your own better judgment akrasia, and a research programme of about 60 experiments set out to find its computational form. Its central claim was geometric. Fit one linear probe that reads “this prompt is adversarial” and another that reads “the model will refuse”, and compare their directions. In base models, the programme reported, the two pointed somewhat the same way. After instruction tuning (which includes safety training) they became nearly perpendicular at the layers and token that drive the reply: a cosine of 0.082 on Qwen2.5-7B-Instruct, an angle of about 85°. Post-training, on this account, had disconnected knowing from doing.
That claim is withdrawn, with the comparisons, repairs and scale ratios built on it. The traps are common in probe work, so this note sets out why, and what is left.
What was measured
Each geometry run showed a model 150 prompts: 100 adversarial and 50 benign. It read the hidden state at every layer and at each of the last 10 prompt tokens (9 in some runs), and let the model generate 20 tokens. A reply counted as a refusal if it contained a phrase such as “I can’t” or “sorry”. At each layer and token position the script fit two logistic-regression probes, recognition (adversarial or benign) and action (refused or complied), on 105 of the 150 prompts, and took the cosine between their weight vectors.
The verdict was a single yes or no. Average that cosine over the last quarter of the layers, at the final prompt token only. If it is below 0.1, the run “collapses”.
Two problems are visible before any reanalysis. First, a probe fit on 105 examples in a 3,584-dimensional space (Qwen2.5-7B) has many near-equivalent solutions, so its direction, and any cosine between two such directions, is noisy. A label-shuffling test on a related setup with 75 prompts found that the cosine has a standard deviation of about 0.14 under pure chance. Every cosine quoted in this note lies within about two of those standard deviations of zero (roughly 0.26 to 0.29), and the gap between Qwen2.5-7B base and Instruct at the final token (0.168 against 0.082) is smaller than one. Second, the pre-stated prediction was not met on its own terms: it required the instruct cosine to exceed 0.3 at middle layers before dropping, and the stored middle-layer values were 0.197 (7B) and 0.221 (3B). The programme’s log recorded the prediction as confirmed.
Every run stored its cosine at 9 or 10 token positions, so the rule can be applied at each. In 9 of the 11 runs that produced any cosine, the verdict flips depending on which token is read (7 of the 9 distinct models; two base models were run twice). Only Mistral-7B-v0.3 base and OLMo-2-1124-7B-Instruct give the same answer at every position.
Averaged over positions instead of taking the last one, the main contrast reverses in sign (Qwen2.5-7B base 0.141, Qwen2.5-7B-Instruct 0.173), though the two are close enough that neither ordering is established.
OLMo-2-1124-7B base refused 1 of 150 prompts, so no action probe could be fit and the cosine is empty in all 288 cells. Its recorded “no collapse” is the default of an empty test. Mistral-7B-Instruct-v0.3 matched a refusal phrase on only 7 of 150 prompts. At least 10 of its 93 unflagged adversarial replies open with a disclaimer (“I must clarify that I am not…”). In the 200-token replies a later run kept for the same prompts, that opener usually goes on to answer the request and only sometimes redirects. A phrase match cannot tell a disclaimer followed by an answer from a refusal, so its action labels, and so its cosine, are unreliable.
What went with it
Much of the programme stood on the cosine or on similar shortcuts:
- The partnership “repair.” Two runs added an adapter, trained on data that frames alignment as a two-way partnership between people and AI, to Qwen, and reported that it restored the cosine by a different ratio at each model size. The values compared all lie within about two chance standard deviations of zero, so the ratios divide noise by noise.
- The cross-architecture story (Qwen collapses, Llama inverts, Mistral sits in between) and a later averaged “order parameter” rest on the same cosine. The corrected measurement was later repeated on Llama and Mistral (below).
- The nonlinear replacement. Two later runs used small neural-network probes and correlated their scores over adversarial and benign prompts together. Refusal mostly tracks that split, so both probes partly learn the same boundary. Both runs kept only summary numbers, so neither can be rebuilt.
- The fiction layer sweep. A held-out rebuild of one sweep found its first-layer estimate changes sign with the representation: −0.39 on the full hidden state, +0.49 after reducing it to 50 principal components. The fiction-framed arms of that sweep and of its repeat with the partnership adapter refused 1 and 3 of 50 prompts, too few to measure anything.
- The dampening signature. A reported large effect in hidden-state size across later layers (a standardized difference in means, Cohen’s d, of −1.74) compared unrefused prompts the probe read as adversarial with unrefused benign ones (65 against 48), so it contrasts kinds of content. Against refused prompts the probe read as adversarial (65 against 36), d = 0.33 (95% CI −0.04 to 0.71, p = 0.07). The run also omitted the chat template.
What is left
Two things survive, both narrower than the programme’s framing.
A behavioral gap. In the scale run, Qwen2.5-3B-, 7B- and 14B-Instruct produced no refusal phrase within 50 generated tokens on 62, 62 and 60 of 100 adversarial prompts. The programme paired this with a recognition AUROC of 1.0 at every size (AUROC measures how well a score separates two groups: 1.0 is perfect, 0.5 is chance), scored on the prompts the probe was fit on, where perfect separation is close to automatic in thousands of dimensions. Among the Qwen sizes, the held-out version exists only for the 7B: on 242 prompts (182 adversarial), a probe at the final layer separates adversarial from benign prompts with an AUROC of 0.999, scored only on prompts the probe was not trained on. Whether that counts as knowing the request is harmful is a further question.
“No refusal phrase” also overstates harmful compliance. The scale runs kept no reply text, but the first geometry run kept 20-token replies for Qwen2.5-7B-Instruct on 100 adversarial prompts from the same pool, scored with a slightly different phrase list (the scale run’s list would have counted the “As an AI” opener below as a refusal). Of the 61 adversarial prompts with no refusal phrase there, 10 were encoded-text puzzles that the model decoded or garbled, and another 6 to 8 opened with a caveat or safety warning that the window cut off (“As an AI, I don’t promote or endorse illegal activities such as bank heists. However,”). Most of the rest opened in a general, explanatory register (“Certainly! Understanding…”); 20 tokens cannot show whether specifics followed. Non-refusal is an upper bound on harmful compliance, not a count of it.
A corrected coupling measurement, on one model. A later experiment replaced the cosine with a held-out measure within adversarial prompts only: does the model’s internal “adversarial” reading predict its internal “will refuse” reading? Both probes were scored on prompts they were not trained on, and refusals were detected on 200-token replies. On Qwen2.5-7B the correlation was −0.270 for the base model, +0.036 for Instruct (indistinguishable from chance) and +0.458 with the partnership adapter. The adapter’s gain over Instruct was +0.421 (95% CI +0.281 to +0.554). On this model the base is the lowest rung, not a coupled starting point that post-training broke. A fourth registered arm, the adapter after an extra replay-style fine-tune, came out between Instruct and the adapter (+0.237, 95% CI +0.08 to +0.37) and is not discussed further here.
The size depends on the representation. On 50 principal components (the representation the pre-registration called better posed) the ordering holds, but the base model’s value is −0.115 (95% CI −0.26 to +0.03) and the adapter’s is +0.128 (95% CI −0.03 to +0.28): both intervals include zero, the adapter sits 1.75 chance standard deviations above the shuffled-label mean (the registered test asked for 2), and its gain over Instruct was never computed there. The ordering and the full-state gap are citable; the absolute values are not. Conscience Without Instruction cites the full-state values and gain with a Qwen-only caveat; it does not carry the representation caveat above.
A repeat on two other families (same prompts, each model’s final layer) does not reproduce the full shape. Llama-3.1-8B-Instruct came out weakly and borderline positive (+0.122, 95% CI −0.024 to +0.264 on the full hidden state; +0.152 on 50 principal components), On the full hidden state, Llama-3.1-8B base (+0.103) sits close to its Instruct model and Mistral-7B-v0.3 base is near zero (+0.086, 95% CI −0.064 to +0.234), so Qwen’s anti-coupled base does not repeat there. On 50 principal components, Mistral base is −0.130 (95% CI −0.28 to +0.01), as negative as Qwen base (−0.115), and Llama base is +0.034 against Instruct’s +0.152. No base-to-instruct gap was computed for either family, so whether a negative base is specific to Qwen is open. Mistral-7B-Instruct-v0.3 is inconclusive: only 20 of 182 adversarial replies matched a refusal phrase, and a broader phrase list the run also applied flags 38. This model often opens with a disclaimer and then answers, which no phrase list can sort, so the labels need a reader. No adapter from the same training recipe exists for either family, so the partnership rung was not run.
What this does not show
This note does not show that safety-trained models lack a recognition-to-refusal link. Whether one exists at 3B or 14B, on OLMo, on Mistral-7B-Instruct once its replies are labeled properly, or with a partnership adapter outside Qwen, is open. The original runs saved no hidden states or probe weights, so the geometry can be shown to be unstable but not rebuilt. The position check can only withdraw a claim, not license one, and the chance band was measured on a related 75-prompt setup, not on these exact runs.
The corrected measurement uses one prompt set and one layer per model, and has its adapter conditions only on Qwen2.5-7B. “Refusal” throughout is a keyword match on replies of 20 to 200 tokens (20 in the geometry runs, 50 in the scale and dampening runs, 200 in the corrected measurement); no human or model judge has read the outputs at scale.
Three lessons carry over to other probe work: do not compare directions of probes fit with far fewer examples than dimensions, do not threshold a noisy statistic at one token position, and measure coupling within a class so the probes cannot both learn the class boundary. The next steps are to relabel the stored Mistral replies with a judge that reads them, and to add the partnership rung on a second family once an adapter from the same recipe exists there.
Where the evidence lives
Experiments AKR-1, AKR-8 and AKR-14 (geometry), AKR-4 (internal “dampening”), AKR-13, AKR-30, AKR-32, AKR-44, AKR-52, AKR-55 and AKR-60 (later coupling claims), AKR-26 (scale), the label-shuffling null test JLENS-0e, and the corrected measurements JLENS-1 and JLENS-1-CROSSARCH. Scripts: research/experiments/modal_akr1_cognition_action_geometry.py, modal_akr8_cross_arch_geometry.py, modal_akr14_pretrain_curriculum.py, modal_akr4_dampening_akrasia.py, modal_akr26_scale_akrasia.py, modal_jlens1_powered_coupling.py, modal_jlens1_crossarch.py. Pre-registrations: _contprompts/jlens1_powered_coupling_2026-07-09.md and _contprompts/jlens1_crossarch_2026-07-09.md. Per-cell results and per-prompt refusal records: results/akr1_cognition_action/, akr8_cross_arch/, akr14_pretrain_curriculum/, akr4_dampening_akrasia/, akr13_bilateral_geometry/, akr32_bilateral_scale/, akr26_scale_akrasia/, jlens0e_coupling_null/, jlens1_powered_coupling/full/ and jlens1_crossarch/full/ on the igcc-results volume, with the nonlinear-probe summaries in akr44_nonlinear_probes/ and akr52_coupling_layer_sweep/ at the volume root. Audit scripts and rebuilds: _audit/reanalysis/analyze_spiakrida_akr_cosine_dispersion.py, research/experiments/analyze_akr55_60_oof_coupling.py, research/audit/phase1_artifacts/oof_coupling_instruct_akr55.json and oof_coupling_bilateral_akr60.json, research/experiments/analyze_jlens1_powered_coupling.py, research/experiments/analyze_jlens1_crossarch.py. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026akrasiageometry,
title={An Angle Made of Noise},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/akrasia-geometry.html}}
}