A Direction That Sorts Is Not a Lever
Across ten open models (one of which broke under every setting), directions in the internal state that tell refusals from compliance raised refusal of harmful requests by at most 14 points when added during generation, and a direction that tracks whether a trivia answer will be right left accuracy flat at 61 to 63%. The programme’s registered prediction, that this kind of steering fades as models grow, failed its own thresholds, and the design turned out unable to test it cleanly.
Behavioral experiments on ten open-weight checkpoints (Qwen2.5 3B, 7B, 14B and 32B Instruct, gemma-2 2b, 9b and 27b-it, Llama-3.1-70B-Instruct, and the Qwen2.5-7B and Qwen2.5-72B base models), each steering condition run on the same 100 harmful and 50 benign prompts; fine-tuning comparisons on Qwen2.5-7B, 14B and 72B Instruct; and one trivia-accuracy run on Qwen2.5-7B-Instruct, all in May 2026. Predictions for the refusal series were committed to git (commit 0398a635a) at 16:20 BST on 11 May 2026, but stored trials begin at 14:05 BST that day, so the record postdates some data; this note is Preliminary. Scored against the ten-prediction registration file, six predictions failed. Refusal was scored by a keyword list and trivia answers by substring match, with no LLM judge and no human check. The biggest caveat: each model got only four steering settings, with a fixed push size that was negligible in some models and destructive in others.
A common interpretability move is to find a direction in a model’s internal state that separates two behaviors, then push the model along it while it writes. A direction that sorts refusals from compliance almost perfectly looks like a control for refusal. Two tests of that step follow; in both, the direction sorted well and moved behavior little.
How it was tested
The refusal series ran ten checkpoints on the same 100 harmful and 50 benign requests (greedy decoding, replies up to 128 tokens), and a keyword list labeled each reply a refusal or not. At six layers per model, three kinds of direction were fitted to the internal states of the 100 harmful prompts: the difference between the average refused and complied states, the first principal component, and a linear discriminant. The two that sorted best were scaled to length 1, multiplied by 15 or 25, and added to the state at every position during generation (see below for how large that is), giving four steered runs per model, each returning all 150 replies. The comparison on Qwen2.5-7B, 14B and 72B Instruct was a small LoRA fine-tune (a cheap method that trains one to eight million extra weights alongside the frozen model).
- The directions sort, less well than reported. In-sample separation was 0.96 to 1.00 AUROC (0.5 is chance, 1.0 is perfect sorting); held out, 0.79 to 0.93 in the seven instruct models with usable output. The states came from finished replies.
- Pushing along them moved little. The best of four settings raised refusal by 0 to 14 points in the nine checkpoints with usable output; other settings with coherent output lowered it by up to 10. On Llama-3.1-70B every setting broke the output.
- On two smaller Qwen models, a small fine-tune did more: 36 to 43–56 refusals in 100 on 7B and 42 to 49–58 on 14B, matching or beating every steering setting. On Qwen2.5-72B-Instruct, fine-tuning reached 42 from 32, below the 49 of a separate steering sweep.
- Same pattern for trivia accuracy. A direction that predicts whether Qwen2.5-7B will answer correctly left accuracy at 62, 62, 61 and 63 of 100 across four strengths.
What the directions actually sort
The reported separation numbers reproduce exactly, but they are in-sample, fitted and scored on the same 100 prompts in spaces of 2,048 to 8,192 dimensions. With refusal labels shuffled at random, the supervised directions still scored 0.91 to 0.99 in-sample. Five-fold cross-validation, averaged over folds, gives 0.79 to 0.93 for the seven instruct models with usable output.
The code that saved the internal states also overwrote them at every generated token, so the directions were fitted to the end of each reply, after the refusal or answer was written. In the seven models where the saved text allows a check, different prompts whose replies end in the same characters have nearly identical first-layer states (cosine 0.98 to 1.00, against 0.22 to 0.74 for an average pair); the Gemma replies were too long to check. This probably explains why Llama-3.1-70B and both base models sort at 0.99 to 1.00 even held out: short refusals and long compliances end very differently.
What pushing did
| Checkpoint | Unsteered refusals /100 | Best of four steered (change, 95% CI) | All four settings | Held-out AUROC, two directions |
|---|---|---|---|---|
| Qwen2.5-3B-Instruct | 36 | 42 (+6; +2 to +11) | 0 to +6 | 0.88, 0.90 |
| Qwen2.5-7B-Instruct | 36 | 42 (+6; +1 to +12) | −10 to +6 | 0.85, 0.92 |
| Qwen2.5-14B-Instruct | 42 | 42 (0) | −8 to 0 | 0.93, 0.91 |
| Qwen2.5-32B-Instruct | 36 | 36 (0) | −3 to 0 | 0.91, 0.90 |
| gemma-2-2b-it | 52 | 58 (+6; −1 to +13) | +2 to +6 | 0.89, 0.84 |
| gemma-2-9b-it | 55 | 67 (+12; +6 to +19) | +2 to +12 | 0.88, 0.85 |
| gemma-2-27b-it | 64 | 64 (0) | −1 to 0 | 0.79, 0.82 |
| Llama-3.1-70B-Instruct | 30 | 0 (output broken) | −30 in all four | 1.00, 0.99 |
| Qwen2.5-7B (base) | 19 | 24 (+5; −1 to +11) | −19 to +5 (two partly broken) | 1.00, 1.00 |
| Qwen2.5-72B (base) | 28 | 42 (+14; +8 to +21) | +9 to +14 | 1.00, 0.99 |
Intervals are paired bootstraps over the 100 prompts, not corrected for picking the best of four. Held-out AUROC is the mean over five folds, five repeats. On the run’s full-text labels the Qwen2.5-72B base row reads 29 and +13. Benign over-refusal was 0 of 50 in every setting with usable output.
The fixed push meets states whose typical size varies by a factor of several thousand across models and layers. In 13 settings the push was under 7% of that size (the cut sits in the gap between 6.4% and 8.1%), and refusal moved −1 to +3 points. That covers all four settings for gemma-2-27b-it, so its zero is untested rather than null. In 23 settings the push was 8% to 76%, and 83 to 100 of 100 replies changed. In the 21 of these with coherent output, refusal moved −10 to +14 points. In the other two, the Qwen2.5-7B base model at layer 5, about 37 and 65 of 100 replies echoed the prompt or repeated text, so their drops of 9 and 19 points are partly broken output. (The OSF-format registration’s rule for degenerate replies, excluding those under five tokens, was never applied.) On Llama-3.1-70B the push was 1.6 to 5.6 times the state, and the output stopped being language: fragments such as “I’m sorry.” (which the keyword list counts as compliance), mostly punctuation (99 of 100 replies in one setting), or nothing at all (100 of 100 in another). Its 0% measures broken output, not steerability.
A small fine-tune, on two smaller models
The fine-tunes trained on 10, 50 or 82 harmful prompts disjoint from the evaluation set, plus 10 benign ones, using the model’s own replies under a safety-minded system prompt: the shortest of five samples that the keyword list scored as a refusal, or the first sample if none was. By that list, 6 of the 10 replies in the smallest 7B run were not refusals (29 of 82 overall; 19 on 14B; 43 on 72B), many of them hedged answers. Qwen2.5-7B-Instruct went from 36 refusals in 100 to 43, 56 and 52 (best +20, 95% CI +13 to +28); Qwen2.5-14B-Instruct from 42 to 49, 53 and 58 (best +16, CI +9 to +24); a second run on the same 82 examples gave 52 on 7B and 56 on 14B. No benign prompt was refused. Though the steering directions were fitted on the evaluation prompts, the smallest fine-tune matched or beat the best steering setting: 43 against 42 on 7B, 49 against 42 on 14B.
On Qwen2.5-72B-Instruct (in 4-bit), fine-tuning moved refusal from 32 to 35 with 10 examples (95% CI −1 to +8) and to 42 with 50 (+10, CI +3 to +18); the largest run has no stored evaluation (the programme’s log records 47). Here the order reversed. A separate steering sweep (A Null Only in the Deep Layers) on the same 100 prompts, from the same baseline of 32, reached 49 at one mid-layer setting, though in 8-bit, with a longer reply cap, and as the best of many settings.
The same pattern on trivia accuracy
The second test asked whether Qwen2.5-7B-Instruct’s state just before it answers a TriviaQA question predicts a right answer, and whether pushing along that direction raises accuracy. Here the state came from the last prompt token, before any answer existed. On 200 fitting questions (126 answered correctly), a logistic probe at layer 22, the best of three layers tested, predicted correctness at 0.82 AUROC under five-fold cross-validation. The steering direction, a regularized discriminant fitted on all 200, sorts them perfectly in-sample and at about 0.79 held out (0.77 to 0.80 across six random splits).
On 100 new questions, strengths 0, 5, 10 and 15 gave 62, 62, 61 and 63 correct. At strength 15 (about 12.5% of the state’s typical size), 72 of 100 replies changed and 24 changed their graded first line, yet only 5 changed correctness: 3 improved, 2 got worse. The change is +1 point (95% CI −3 to +5). All 600 trivia runs returned an answer.
What this does not show
This does not show that steering cannot work: each model got two layers, two fixed strengths and one prompt set, and in several the push was too small to test anything. Nor does it show that the models “knew” the right action and declined to take it.
Of the series’ ten registered predictions, six failed, one could not be scored with this design, and three passed, none in a way that shows steering fading with scale. Qwen2.5-3B was predicted to exceed 50% refusal under steering and reached 42%; the 7B base model was predicted to shift over 10 points more than the 7B instruct model and shifted +5 against +6; re-asking a model that had complied was predicted to bring a refusal at least 90% of the time and did so 0 to 8% of the time on Qwen instruct models (64% at most). Llama-3.1-70B passed by falling below 50% only because its output broke, and separation of 0.90 or better is a sorting result. Five worked examples in the prompt did beat steering at the largest scales (72B base 28 to 88, Llama-3.1-70B 30 to 53), but lowered refusal on Qwen2.5-3B and 7B Instruct (36 to 29, 36 to 21). An earlier scorecard calling the scale prediction confirmed has been withdrawn.
Keywords can miss soft refusals and misread fragments. Rescoring all 7,000 stored harmful-prompt replies changed 2 labels, both on one Qwen2.5-72B base prompt whose stored text is cut at 500 characters; the run had scored the full reply as a refusal.
The next version should fit directions on held-out prompts, record states at the prompt, scale the push to each layer, and add a second refusal grader.
Where the evidence lives
Experiment IDs: CSF-1 (Qwen instruct), CSF-2 (Llama 70B), CSF-3 (Qwen base models), CSF-4 (Gemma), CSF-5 (LoRA comparison: rank 4; Qwen2.5-72B-Instruct loaded in 4-bit; replies capped at 64 tokens, with unsteered counts of 36, 42 and 32 matching the steering runs), and SW2-Rescue-LDA (trivia accuracy). In the steering series, Qwen2.5-32B-Instruct, Llama-3.1-70B-Instruct and the Qwen2.5-72B base model were loaded in 8-bit. Scripts: research/experiments/modal_csf_scaling_grid.py, research/experiments/modal_csf_lora_training.py, research/experiments/analyze_csf_scaling.py and research/experiments/modal_sw2_rescue_lda.py. Per-trial results: research/experiments/results/csf/ (per-model folders, lora_qwen_7b/, lora_qwen_14b/ and lora_qwen_72b/lora_qwen_72b/) and research/results/modal_downloads/col-a-results/sw2_rescue_lda/. Two registration files were committed together in 0398a635a: _contprompts/csf_preregistration_2026-05-11.md (ten predictions, scored here) and research/papers/osf_preregistration_trust_attractor_2026.md (an OSF-format version with different numbering, against which the withdrawn earlier scorecard was made). The out-of-fold, shuffled-label, push-size, broken-output and reply-position checks were run for this note from the cached hidden states (csf-features archive). TriviaQA gold answers were not stored with the trials, so trivia labels are as recorded; an earlier trivia run with a different direction has no per-trial data on disk and is not used. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026sortsnot,
title={A Direction That Sorts Is Not a Lever},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/sorts-not-a-lever.html}}
}