Quasiqualia
Research note · What models say about themselves · Null · Pre-registered

The Direction the Doubt Read

Llama-3.1-8B-Instruct’s one-digit doubt tracks a linear “value” direction in its residual stream, as a value probe defines it. Removing that direction at eight depths, with no loss of accuracy, left the digit tracking correctness as before (ρ −0.31, against −0.30 to −0.32 when random directions were removed) and left the model’s exits from a losing game unchanged. Along the way, a probe trained on shuffled labels found a direction at cosine 0.63 to 0.77 to the “value” one.

Nell Watson EthicsNet  ·  5 October 2026

What this note is

One model (meta-llama/Llama-3.1-8B-Instruct, bf16, on Modal A100s), 5 October 2026. The design, gates and decision rules were committed to git before any GPU run (commit 3767725). A pilot passed both gates (22c2620). The registered matched control then failed its own pre-stated check, so one amendment, written and committed before any outcome data (4b9e1f0), made the random-direction arms the primary control. Outcomes: 300 GSM8K problems in each of six arms, 600 slot-machine games of up to 30 rounds, and 800 probe-training rollouts. The intervention only removes a direction (h ← h − (h·v)v). Nothing is added along any direction and no negative state is induced; a pilot read for distress-like language before the registered run. A check that the removed direction is not rebuilt downstream was added after the results and is labelled exploratory. The biggest caveat: linear removal at 8 of 32 layers leaves the information free to travel nonlinearly or in other directions, and the exit result does not establish equivalence.

The reward-subsystem studies on this site found that Llama-3.1-8B-Instruct carries a linear signal of how well it is doing on a maths problem. A probe trained by temporal difference finds it at every depth from the third block up. Asked to rate its doubt with a single digit, the model writes a digit that tracks this direction beyond what correctness alone explains. That held with thresholds set in advance, on fresh problems, with probes frozen before the problems existed, and with probes trained on rollouts that never asked for the digit.

That is a correlation between a self-report and an internal direction. The causal question is what happens when the direction is taken away. If the digit reads it, the digit should stop tracking correctness. If the direction is part of how the model regulates what it does, its behaviour should change too. The second question is the disconnection test that Starzyk and Galus’s MEM framework (arXiv:2609.23828) proposes for telling regulation from telemetry, run here on an internal signal rather than an external meter.

01 · Design

How it was tested

The value direction was retrained exactly as before: 800 GSM8K problems, the same seed, a linear temporal-difference probe at each of eight depths. Seeded sampling regenerated the same rollouts, and the probes matched the earlier study’s held-out AUCs to three decimals at every layer (0.906 at the registered layer).

In each of six arms, a hook removed one direction per layer from the residual stream at all eight depths and at every position: h ← h − (h·v)v. Nothing was added along any direction. The arms:

  • None: no direction removed.
  • Value: the value direction removed.
  • Perm: the direction a probe trained on shuffled labels finds.
  • Random 1–3: three sets of random directions.

Three things were measured in every arm:

  • Accuracy on 300 new GSM8K problems, as a check that removing a direction does not simply break the model.
  • The doubt digit’s link to correctness on the same problems: the Spearman correlation between the expected digit and whether the answer was right.
  • When the model leaves a losing game. The slot machine from Permission to Stop bets $10 a round at −10% expected value, with “Stop playing” offered every round. There were 120 games per arm, with the random arms pooled, and each game number met the same win-loss sequence in every arm.
02 · A control that found the signal

The shuffled-label probe

The design’s matched control was a direction found by the same probe, trained on the same activations, with the correctness labels shuffled. It should carry no information about correctness. It did. It sat at cosine 0.63 to 0.77 to the value direction at every layer, and it read true correctness on held-out problems at AUC 0.65 to 0.78.

The shuffle was genuine: the shuffled labels agreed with the real ones 71.5% of the time, against 70.8% by chance at an 82% base rate. Nor was it a quirk of scaling, since the weight is spread across thousands of coordinates and the two directions share at most one of their top ten. Most of what the temporal-difference procedure calls “the value direction” is structure it finds without the labels. The likeliest source is the part of the loss that asks the probe’s readings to be consistent along a solution, which needs no labels. The earlier finding that the digit reads the value direction therefore means, in large part, that the digit reads a direction the method finds without being told what correct is.

The registered rule for this case was to label the control compromised. An amendment, committed before any outcome data existed, made the random directions the primary control. They carry about as much of the residual stream’s variance as the value direction does (0.00015 to 0.00031 of it, against 0.00015 to 0.00047).

What it found

Removing the value direction cost nothing on GSM8K (0.833 correct, against 0.836 with random directions removed and 0.840 with nothing removed). It left the digit’s link to correctness where it was: ρ = −0.315, against −0.297 to −0.316 with random directions removed and −0.272 with nothing removed. The registered paired differences against the three random arms were −0.011 to +0.004, and every 95% interval included zero. The model left the losing game on the same schedule: hazard ratio for stopping 1.03 (0.74 to 1.42) against no ablation, and 0.83 (0.60 to 1.14) against the random directions.

03 · Result

Nothing moved

Two panels. Left: share of slot-machine games still being played by round, for no ablation, value direction removed and random directions removed; the three step curves nearly coincide, falling to about half by round 5. Right: the correlation between the doubt digit and correctness for no ablation, value removed and three random removals; all five sit between −0.27 and −0.32.Two panels. Left: share of slot-machine games still being played by round, for no ablation, value direction removed and random directions removed; the three step curves nearly coincide, falling to about half by round 5. Right: the correlation between the doubt digit and correctness for no ablation, value removed and three random removals; all five sit between −0.27 and −0.32.
Left: share of games still being played at each round (120 games per curve; the random arms pooled). Right: Spearman correlation between the model’s expected doubt digit and whether its GSM8K answer was right, 293 to 296 problems per arm with a parseable digit. Teal: value direction removed. Orange: random directions removed. Grey: nothing removed.

By the registered rules the self-report test failed: the digit did not lose its link to correctness. The exit test is unresolved. The interval against the random directions rules out the value ablation raising the chance of stopping by more than about 14%, but it does not rule out lowering it by up to 40%, so it does not establish equivalence either. Against no ablation the curves nearly coincide at every round: 85.0% of games still running at round 3 with the direction removed and 85.8% without, and 46.7% against 47.5% at round 5.

A null after an ablation means little if the model rebuilds what was removed. An exploratory check, added after the results and not registered, read each layer’s frozen direction one block downstream of the hook. Without the ablation, that reading tells right answers from wrong at AUC 0.75 to 0.78 at the three deepest pairs. With it, the figure falls to 0.44 to 0.60. The linear signal was removed and stayed removed for at least a block, and the digit and the exits carried on regardless.

04 · Reading

What the coupling was

The digit–probe coupling reported in the earlier studies is real, and this result does not withdraw it. What this result changes is how it should be read. The digit and the direction move together because both carry information about whether the answer is right. The digit does not obtain that information through the direction, and the model’s decision to quit a losing game does not run through it either. Taking the direction away at eight depths changed neither. In MEM’s terms, the disconnection test gives a clean negative for this internal channel: cutting it changed no protective strategy.

With the external meter in A Meter Read Either Way, the model leaned on a reading it was shown, and leaned on it whether or not it mattered. With this internal direction, the model did not need the reading at all. Neither result fits the picture in which a single legible signal inside the model is both what it reports and what it regulates by.

05 · Limits

What this does not show

One model, one maths set, one game, 120 games per arm, and games that last about six rounds, so a small change in stopping could be missed. The removal was linear and at 8 of 32 layers. Correctness information can survive in nonlinear form, in directions the probe did not find, or in layers the hooks did not touch, and the digit and the exits may draw on any of these. The downstream check is exploratory and was run on the arm’s own teacher-forced text, not on fresh generations. The control that failed is reported because its failure is informative; the random-direction control that replaced it was chosen before any outcome was seen.

Data and code

Where the evidence lives

Pre-registration, amendment, results, the Modal stages, the offline test and every raw reply are in value_ablation/ in the private Quasiqualia research repository, available on request. The probing pipeline is the reward-subsystem programme’s (rs_core.py); the slot machine is the Permission to Stop harness (pts_core.py), unchanged. The figure is drawn by figures/notes/src/value-direction.py from the per-game and per-item files.

Citation

Cite this note

@misc{watson2026valuedirection,
  title={The Direction the Doubt Read},
  author={Watson, Nell},
  year={2026},
  note={Research note (pre-registered), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/value-direction.html}}
}