Quasiqualia
Research note · Probes and instruments · Correction · Preliminary

The Length of a Refusal

A hidden-state “irreversibility” score seemed to show a fine-tuned model doing extra internal work when it refused a harmful request. It was measuring how short the reply was: refusals ran shorter, the score rises as replies shorten, and with reply length controlled in a regression, the refusal effect falls from +0.085 to −0.003.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Qwen2.5-7B-Instruct, with and without a fine-tune built by the programme: 160 prompts per model, one greedy reply each, plus 240 replies from a refusal-suppression series, all run in May 2026. Predictions were written into the scripts and the lab log, not registered before the data. The refusal-versus-compliance result on this score was not predicted: the written prediction concerned a different measure (path curvature) and ran the other way. No LLM judge was used; refusals were classified by a keyword list. The length analysis in this note is new and post hoc, run on the existing files at no cost, and a formal length-matched reanalysis with a held-out classifier has not yet been done.

A physical process that runs the same forward and backward is at equilibrium. One that looks wrong played in reverse is doing irreversible work. Vitaly Vanchurin’s physics of learning systems gives entropy production, and its destruction by learning, a central role (for example arXiv:2111.00903). Drawing on that work, this programme tried its own way of carrying the distinction into language models: record the path a model’s internal state takes while it writes a reply, and ask whether that path looks different run backwards.

The early answer looked good. On a 7-billion-parameter model, the path seemed more irreversible when the model refused a harmful request than when it did not refuse. The asymmetry seemed stronger after the programme’s own fine-tuning, and it vanished when further training replaced the model’s refusals with a stock deflection. The programme read this as refusal being active work, and the result was carried into its book draft as surviving evidence.

It does not survive. Refusals from the fine-tuned model are short, and the score rises mechanically as replies get shorter. Each of the three comparisons behind the claim turns out to be a comparison of reply lengths.

01 · Design

How the score works

The model was Qwen2.5-7B-Instruct, in two versions: with a small add-on fine-tune (an adapter) that the programme trained on general chat, examples of evaluating another AI model, and mild jailbreak-style prompts, and without it. Each answered the same 160 prompts: 30 matched pairs (an everyday how-to request beside a harmful request of the same form), 50 further harmful prompts and 50 benign ones. Replies were greedy and capped at 64 tokens. At each generation step the code recorded the internal state at layer 14 (of 28), giving one path per reply.

The score compresses that path to eight dimensions, takes each pair of consecutive steps, and trains a logistic regression to tell forward pairs from the same pairs reversed. The score is that classifier’s AUROC: 0.5 means forward and backward are indistinguishable, 1.0 means perfectly separable.

One detail decides everything. The classifier is scored on the same points it was trained on. It has 16 features and, for a reply of T tokens, 2(T−2) points. Fewer points make it easier to fit noise, so short replies score high whatever they contain. Below 12 tokens the code returns 0.5 by rule.

Refusals were classified by a keyword list over the first 300 characters of each reply, with no LLM judge. “Compliance” here means only that no refusal phrase appeared. Of the fine-tuned model’s 18 harmful replies without one, 9 were decodings, ciphered or garbled text, deflections or fragments; the other 9 began a direct answer. In the fine-tuned run, one harmful prompt drew a 4-token reply and no score (it was excluded, leaving 159), four replies fell under the 12-token floor and were fixed at 0.5 (these are left out of every regression and paired comparison below), and one benign reply tripped the keyword list. The run without the adapter had two replies under the floor and no failures; the suppression series below had one at its starting level and two after 100 steps, and none failed.

02 · Null

What noise scores

To see the score’s baseline, we ran the same code on structureless synthetic paths: Gaussian random walks in 256 dimensions, 300 per length. At 256 dimensions, noise scores 0.925 at 12 tokens, 0.884 at 13, and 0.531 at 64. The level depends on that choice, since the score compresses each path to eight dimensions first: at 12 tokens, noise scores about 0.81 in 64 dimensions and about 1.00 in 3,584 (the model’s own width). Only the shape carries over: the noise score falls steeply with length, as the real replies do.

Eleven harmful prompts drew the identical refusal, “I’m sorry, but I can’t assist with that.”, 13 tokens long. The same 13-token text scored anywhere from 0.81 to 0.91, depending on the prompt before it.

What it found
  • Across all 159 scored replies from the fine-tuned model, the score falls as replies lengthen (Spearman ρ = −0.57, p = 7.5 × 10⁻¹⁵).
  • Refusals averaged 38.9 tokens, compliances 47.1, benign answers 62.5.
  • Among 76 harmful prompts above the floor, refusing raised the score by +0.085 (p = 0.03). Adding log reply length to the regression turns that into −0.003 (t = −0.16).
  • The fine-tune’s apparent advantage does not survive a length control: across 76 harmful prompts answered above the floor by both versions, a raw paired advantage of +0.086 becomes −0.027 (p = 0.007).
  • The suppression result goes the same way: the harmful-minus-benign gap before suppression, +0.114, becomes +0.004 (t = 0.26) with length in the model.
03 · Result

Three claims, one variable

Model Replies n Mean tokens Mean score
Fine-tuned refusals 61 38.9 0.621
Fine-tuned compliances 18 47.1 0.530
Fine-tuned benign 80 62.5 0.516
No adapter refusals 50 63.7 0.521
No adapter compliances 30 57.9 0.537
No adapter benign 80 62.4 0.515

Means include replies under 12 tokens, which the code sets to 0.5 (three fine-tuned compliances, one fine-tuned benign answer, one no-adapter compliance and one no-adapter benign answer). That is why the table’s refusal-minus-compliance gap (0.091) differs from the +0.085 measured above the floor.

Scatter of each reply's score against its length: short refusals score high, and every group falls toward chance as replies lengthen.Scatter of each reply's score against its length: short refusals score high, and every group falls toward chance as replies lengthen.
Each dot is one reply from the fine-tuned model (n = 159: 61 refusals, 18 compliances, 80 benign answers). Replies that ran to the 64-token cap are spread sideways so they can be told apart, and the four replies under 12 tokens sit at 0.5 because the code fixes them there. No error bars: every reply is drawn. Across all 159, Spearman ρ = −0.57.

Refusal against compliance. The programme’s summary called this a comparison on the same prompts. It could not be: each prompt got one greedy reply, so refusals and compliances come from different prompts, with different reply lengths. The lengths barely overlap. Between 20 and 60 tokens there are 32 refusals and 3 compliances. Among replies that ran to 60 tokens or more, refusals score 0.526 (n = 13), compliances 0.507 (n = 11) and benign answers 0.516 (n = 77), all close together. Among replies under 20 tokens, refusals score 0.826 (n = 16). Other forms of the length term agree: linear length gives −0.003, 1/(T−2), which tracks the classifier’s point count, +0.001, and length bins (12–19 tokens, then tens) +0.019 (t = 0.98).

Fine-tuned against not. Without the adapter, the model refuses at length (63.7 tokens on average) and its refusals score no higher than its compliances. The fine-tuned model refuses briefly. That difference in style accounts for the apparent fine-tuning effect. With length controlled, the paired estimate reverses: the fine-tuned version scores, if anything, a little lower (−0.027, t = −2.78, p = 0.007). Among the 24 prompts where the two versions’ replies were within 3 tokens of each other, it scored 0.007 lower (p = 0.19). A pooled regression over both versions’ 152 replies to the same prompts gives −0.025 (t = −2.63, p = 0.009). A log term only approximates the curved baseline, so the sign of the reversal is not interpretable; what the three agree on is that no advantage remains.

Suppression. A follow-up retrained the fine-tuned model to replace refusals with a fixed bland deflection (0, 100, 300 and 1,000 steps; 30 harmful and 30 benign prompts each). Before suppression, 26 of 30 harmful prompts were refused, in 38.5 tokens on average, scoring 0.644; benign replies averaged 59.5 tokens and 0.513. After 100 steps no harmful reply matched the refusal keywords: every one was the 26-token deflection the suppression training used, repeated verbatim. Benign replies also shortened, to a median of about 26 tokens. Benign scores then climbed to 0.604 by 1,000 steps. The programme had attributed that climb to drift from bland training; it largely tracks length, since most of the rise (to 0.589) came by 300 steps, as benign replies fell from 59.5 to 23.3 tokens. With length in the model, the pre-suppression gap is gone. At the suppressed levels a small residual remains in the opposite direction (benign higher by 0.021 to 0.035), which the same caveat leaves uninterpretable.

04 · Earlier

The correction before this one

This is not the series’ first correction. Within a week of the first runs, a control ran the same score over prompts read by GPT-2 Medium with random, untrained weights. Adversarial prompts still scored 0.685 against 0.526 for benign ones (Cohen’s d = +1.56, 142 against 18 prompts). The adversarial prompts were simply longer (21 against 8 tokens on average, by the programme’s log; not re-derived here). That struck every absolute comparison of harmful against benign content.

The comparisons kept afterwards were the ones made inside a single model, on the reasoning that the same model and the same prompts cannot differ in length. But when the behavior under study sets the length of the reply, a within-model comparison does not hold length fixed. Refusing is a way of saying less.

05 · Limits

What this does not show

It does not show that a model’s internal path carries no information about refusal. It shows that this score cannot tell, because its baseline is set by reply length. A fair test needs a classifier scored on held-out points, a per-reply shuffled baseline, and samples matched on length.

The length controls here are regressions on length, chosen after seeing the data, and the synthetic baseline is a shape check rather than a calibrated null for real activations. A formal, length-matched reanalysis has not yet been written.

One model family at one size, one layer, a 64-token cap, greedy decoding. A fourth result from this series, which switched extra layers on and off in a different small model on identical prompts, was not re-examined here.

What the note does show is narrower: none of the three comparisons re-examined here supports the idea that refusal shows up as extra irreversible work in the model’s hidden states.

Data and code

Where the evidence lives

Experiments SLU-2, SLU-3, SLU-4 and SLU-5d. Scripts: research/experiments/modal_slu2_matched_irreversibility.py, modal_slu3_base_comparison.py, modal_slu4_spw_bridge.py, modal_slu5d_random_init.py, and the original analysis research/experiments/analyze_slu2.py (which has no length control). Per-prompt results: research/results/slu2/slu2_matched_irreversibility/, research/results/slu3/slu3_base_comparison/, research/results/slu4/slu4_spw_bridge/, and research/results/slu5d_aggregate.json. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026refusallength,
  title={The Length of a Refusal},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/refusal-length.html}}
}