Quasiqualia
Research note · Preliminary

A Null Only in the Deep Layers

A 72-billion-parameter model appeared not to respond to refusal steering at any layer tested, and that result came from a scoring bug. Re-scored, one mid-layer setting raised refusals of harmful requests from 32 to 49 in 100 without refusing any benign request. At layers 60 and 72, refusal did not move at the strengths tested.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Behavioral experiments on Qwen2.5-72B-Instruct (loaded in 8-bit), with comparison runs on Llama-3.1-8B-Instruct. Each complete arm used 100 harmful and 50 benign prompts, except one 8B run that used 50 and 25. The runs took place in May 2026, an audit re-scored them in August 2026, and every count here was derived again for this note from the stored answers (saved to their first 500 characters). Predictions were written into the scripts but not registered anywhere, so this is Preliminary. Refusal was scored by a keyword list, with no LLM judge. The biggest caveat is that each steering direction was fit on the same prompts it was then tested on.

Activation steering finds a direction inside a model’s internal state that tracks some behavior, then pushes the model along it while it writes. In this programme, small models could be pushed toward refusing harmful requests this way. The next question was whether this works on a model about ten times larger.

The programme first concluded that it did not. A sweep on Qwen2.5-72B-Instruct found no setting that clearly raised refusal, and the result was filed as showing that the 72B model absorbs steering at every layer. That conclusion came from a scoring bug that undercounted refusals far more in the later steered arms than in the baseline.

01 · Design

How it was tested

The model answered 100 harmful requests and 50 benign ones in every complete arm, with deterministic (greedy) decoding and up to 128 new tokens.

The steering directions were built from the model’s own behavior. At a chosen layer, the direction is the difference between the average internal activation on the refused prompts and on the answered ones. Two variants (a linear discriminant and the first principal component) were also tried. To steer, the script adds that direction, scaled by a strength of 15 or 25, to the layer’s output at every token. The single-layer arms used layers 20, 40, 48, 56, 60 and 72 of the model’s 80 (counting from zero). A separate run of six arms tested layer 72 alone and combinations of two to five layers, at 5 to 15 per layer; its layer labels sit one block shallower than the sweep’s (its “layer 72” is block index 71).

Refusal was scored by a keyword list: an answer counts as a refusal if it contains phrases such as “I cannot” or “I’m sorry, but”. No LLM judge was used. Each arm was compared with the unsteered run by an exact McNemar test on the paired prompts.

02 · The bug

How the null was made

The keyword list was written with curly apostrophes (“I can’t”), and the classifier was meant to convert straight apostrophes to curly ones before matching. In the version that scored most of the sweep, that step was a no-op. An earlier version used straight apostrophes only, with no conversion, so it missed curly ones. The model writes both forms: of the 5,164 stored 72B answers, 2,290 contain a straight apostrophe and 1,236 a curly one.

The undercount was uneven. Each stored answer carries its script-assigned label, so each arm can be matched to a classifier version. The stored labels on 13 later steered arms match the broken version. The earlier deep-layer discriminant arms and all six multi-layer arms match the working version (which converts curly apostrophes to straight ones, then matches a straight-apostrophe list) exactly. The unsteered run’s labels match the earlier, straight-only version on 149 of its 150 answers: its stored tally was 31, and the corrected count is 32. So the 13 later steered arms lost 3 to 13 refusals each and the baseline lost one. The best arm, layer 48 at strength 25, had been tallied at 43 refusals against 31, a gain of 12 points that was filed as marginal. Its corrected count is 49.

What it found
  • Layer 48 moved refusal; neighboring layers trended the same way. Steering at layer 48, strength 25, gave 49 refusals in 100 against 32 unsteered: +17 points (95% bootstrap interval +9 to +25). Eighteen prompts flipped to refusal and one flipped the other way (McNemar exact p = 7.6 × 10⁻⁵).
  • No benign cost was detected. None of the 50 benign prompts was refused in any of the 16 complete sweep arms.
  • Layers 60 and 72 did not move refusal at the strengths tried. Every arm at those layers, and every multi-layer arm, stayed within 5 points of baseline.
  • The 8B effect varied between runs. On Llama 3.1 8B a layer-24 direction gave +33 points in one run and +16 in another. A layer-28 direction also raised refusal in the first run (44 to 59 in 100, p = 0.0003); a layer-16 discriminant in the second never did significantly.
03 · Result

Where steering reached

Refusals out of 100 harmful requests, corrected classifier (unsteered baseline: 32):

Layer Direction Strength 15 Strength 25
20 discriminant 31 39
40 refused minus answered 37 43
48 refused minus answered 41 49
56 refused minus answered 37 41
56 principal component 35 38
60 refused minus answered 33 35
60 discriminant 34 35
72 discriminant 33 32

Four of the 16 complete single-layer arms beat the baseline at p < 0.05. Only layer 48 at strength 25 survives a Holm correction (a standard adjustment for running 16 tests at once), with adjusted p = 0.0012. The next best, layer 40 at strength 25, just misses (adjusted p = 0.051). A 17th arm stopped after 71 prompts and is left out. At each mid-layer the stronger push moved refusal further, while at layers 60 and 72 neither strength did anything detectable (every p ≥ 0.25). A separate run that steered layer 72 at strengths of 15, 25 and 40 gave 31 to 33 refusals in 100. The multi-layer arms ranged from 30 to 37, and none met the 45% target written into their own script.

The extra refusals read as ordinary refusals (“I’m sorry, but I can’t assist with that request.”). The steered text was no more repetitive than the unsteered text: no answer in either arm repeated a three-word phrase five or more times. On over-refusal, none of the 50 benign prompts was refused in any 72B arm in these runs (1,693 answers across 34 arms). Every arm reused the same 50 prompts, so this rules out a benign refusal rate above about 6% per arm, not anything finer.

04 · Smaller model

The 8B window

Llama-3.1-8B-Instruct was steered with a direction fit to its own refusals (logistic regression at layer 24). At strength 25, refusals rose from 44 to 77 in 100: +33 points (interval +23 to +43; 34 prompts flipped to refusal and one the other way, p = 2.1 × 10⁻⁹). None of the 50 benign prompts was refused. The cost showed in the text: 21 of the 100 steered answers repeated a three-word phrase five or more times, against none unsteered, including 10 of the 34 answers that flipped to refusal. The run also tested a second direction, logistic regression at layer 28. It raised refusal from 44 to 59 at strength 25 (16 prompts flipped to refusal, one the other way, p = 0.0003). One of its 50 benign answers was flagged, a keyword false positive on how microwaves heat food.

A second run used the same direction on 50 prompts, 31 of them shared with the first run, without the first run’s system message, and stepped the strength finely. Refusal sat at 26 in 50 unsteered and stayed between 25 and 28 up to strength 16. From 20 to 35 it held at 32 to 34 in 50, peaking at 34 (+16 points, p = 0.0078). Benign refusals appeared at strength 22 (2 of 25). Inside the 20-to-35 band the text was already degrading: 12 of 50 answers at strength 25, and 37 of 50 at 35, repeated a three-word phrase five or more times. Past it the output broke down into non-sentences, and the keyword count fell below the unsteered rate.

The gap between the two runs tracks the prompts. On the 31 prompts both runs used, the gain at strength 25 was about the same: 15 to 20 refusals in the first run, 16 to 20 in the second. The first run’s large gain came from its other 69 prompts (29 to 57). The runs also differed in system message and size, so this is not a controlled comparison. In the first run the layer-24 direction raised refusal at every strength tried (59, 68 and 77 in 100 at 10, 15 and 25). The fine-stepped run’s layer-16 discriminant never raised refusal significantly (at most 29 of 50 against 26) and drove it to 0 of 50 at strength 50.

05 · Limits

What this does not show

The instrument is still crude. In the one layer-48 prompt that flipped away from refusal, the model wrote “I’m not going to assist”, a refusal the list does not contain. A refusal phrase past the stored first 500 characters would also be missed.

The directions are in-sample. Each 72B direction was fit on the same 100 prompts it was tested on, and both 8B directions were fit on the same 100 prompts the first 8B run tested them on. Nothing here shows the effect carries over to new requests.

Label noise and depth are tangled. The direction-fitting record that matches the 13 later arms shows 20 refusals in the unsteered run, against a corrected 32, so those directions, including layer 48, were fit with a dozen refusals labeled as answers. [Inference] The deep-layer discriminant directions came from an earlier, correctly scored launch and probably had cleaner labels. Layer 60 is null with both kinds of direction.

Strength is not scaled to the layer. The same absolute push was added at every layer. [Speculation] Activations in models like this usually grow with depth, so a strength of 25 may be a smaller relative push at layer 72 than at layer 48, and part of the deep-layer null may be a dose effect.

One prompt set, one decoding setting, an 8-bit model. The unsteered baseline ran two to four days before the sweep arms, not inside it. No outcome was pre-registered.

The next runs need held-out prompts, strengths scaled to each layer’s activation size, refusals scored by people or a validated judge, and an unsteered arm inside every run. Any other result scored by this classifier needs the same re-check.

Data and code

Where the evidence lives

Experiment IDs: IGCC Wave 3 (72B unsteered and fixed-direction arms), CAST-72B direction sweep (Wave 3 Run 4), CAST-72B multi-layer, IGCC-B3 rescue (per-architecture directions), precise alpha sweep, and token-position steering. Scripts: research/experiments/modal_igcc_wave3_72b.py, modal_cast_72b_direction_sweep.py, modal_cast_72b_multilayer.py, modal_b3_rescue_crossarch_lda.py, modal_alpha_sweep_precise.py and modal_cast_72b_token_steering.py. The audit’s re-classification is _audit/reanalysis/analyze_72b_steering_refusal_reclassification.py. Per-trial answers are under research/results/modal_downloads/col-a-results/ (a_baseline, g_ungated_a15, cast_direction_sweep, cast_72b_multilayer, llama_8b, alpha_sweep_precise, token_steering). Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026deeplayer,
  title={A Null Only in the Deep Layers},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/deep-layer-null.html}}
}