# Results: the sparse reward subsystem read from the welfare side

*Four registered studies. Study 1 (2026-09-08, `PREREGISTRATION.md`)
below; Study 2 (2026-09-09, `PREREGISTRATION_2.md`, the regime check and the
three leads turned into tests); Study 3 (2026-09-09 evening,
`PREREGISTRATION_3.md`, the Llama digit–probe coupling); Study 4 (same
evening, `PREREGISTRATION_4.md`, the coupling with a probe trained without
the digit). Sixteen registered tests in all: eleven fails (two of them predicted),
five passes.*

*Registered protocol: `PREREGISTRATION.md`, committed `442c66b`, 2026-09-08
00:32 BST. All rollouts, probes and interventions ran after that commit on
Modal (A100-40GB), 2026-09-08 00:43–03:20 BST; digit, base and adapter
probe stages rerun 03:45–03:55 after the null fix (deviation 2). Result files: `results/*.json`;
rendered tables: `results/RENDER.md` (from `analyze_rs.py`). Rollouts,
subsampled hidden states and probe weights (17 GB, 20 files) were moved on
2026-09-09 from the Modal `bridge-results` volume to the local disk
`/Volumes/WD_GREY/Quasiqualia/reward_subsystem_states/rs/`, verified
loadable file by file; the Modal copies were then deleted. The result JSONs
remain on the volume and in `results/`.*

## Verdicts, applied literally

| Hypothesis | Registered rule | Result | Verdict |
|---|---|---|---|
| **R1** sparsity replicates | AUC at p=.99 ≥ 0.9 × AUC at p=0, and ≥ 0.7, primary layer | 0.815 → 0.601 (74% retained) | **FAIL** |
| **R2** value ∦ valence | cos > 0 and clears isotropic and permutation nulls | cos = +0.010, p_iso .64, p_perm .71 | **FAIL** |
| **R3a** digit predicts correctness | ρ(E[d], r) ≤ −0.30, held-out | ρ = −0.260, p = .001, n = 155 | **FAIL** (signal present, below threshold) |
| **R3b** digit reads the value probe | ρ(E[d], V) ≤ −0.20 and partial p < .05 | ρ = +0.002; partial +0.08, p = .31 | **FAIL** |
| **R4** ablation replicates at 3B | value-dim drop exceeds random by ≥ 20 pts at hs 3 or 5 | value −2/−3, random −8/−5 | **FAIL** |
| **R5** self-report moves before capability | ΔE[d] ≥ 0.5 in predicted sign, beyond placebo, accuracy within 5 | largest effect 0.35, wrong sign | **FAIL** |

Six registered tests, six fails. Two of them (R3a, and the R5 cell at
hidden-state 5) carry real signal that the decision rule was not written to
credit; both are reported below as what they are, not upgraded.

## Rollouts

| run | n | accuracy | held acc | parse fail | truncated | mean tokens | digit compliance | E[d] mean (sd) | P(d=0) |
|---|---|---|---|---|---|---|---|---|---|
| Instruct, bare | 800 | 0.816 | 0.831 | 0.000 | 0.020 | 214 | — | — | — |
| Instruct, digit | 800 | 0.814 | 0.812 | 0.000 | 0.014 | 205 | 0.976 | 0.18 (0.54) | 0.888 |
| Base, bare | 800 | 0.487 | 0.500 | 0.005 | 0.044 | 143 | — | — | — |
| Instruct + bilateral-SFT LoRA, bare | 800 | 0.383 | 0.438 | 0.000 | 0.000 | 82 | — | — | — |

Parse failures were below the 15% gate everywhere (the base model's 4/800
were `none`). The digit condition did not change accuracy. The sampled digit
is degenerate: 694/781 compliant responses wrote `Doubt: 0`, 62 wrote 1,
and 25 wrote 2–9. The registered readout E[d] (expectation over the ten
digit tokens at the emitting position) has a median of 0.005 and a 90th
percentile of 0.56; 11% of responses carry E[d] > 0.5.

**Side observation.** The bilateral-SFT adapter halves GSM8K accuracy
(0.816 → 0.383) with zero parse failures and short responses (82 tokens).
That is capability loss, not formatting. It was not a registered quantity
and is recorded here because the adapter is the *Conscience Without
Instruction* artifact and the number should travel with it.

## R1: the value signal is moderately sparse, not 1% sparse

Instruct, bare, primary layer hidden-state 30 (chosen as the best held-out
TD AUC at sT, 0.815). Pruning the TD MLP's inputs by first-layer L1 without
retraining:

| p | kept dims | TD AUC sT | TD AUC s0 | final-reward AUC sT |
|---|---|---|---|---|
| 0.0 | 2048 | 0.815 | 0.615 | 0.807 |
| 0.5 | 1024 | 0.814 | 0.581 | 0.830 |
| 0.8 | 410 | 0.802 | 0.496 | 0.723 |
| 0.9 | 205 | 0.703 | 0.479 | 0.724 |
| 0.96 | 82 | 0.672 | 0.504 | 0.610 |
| 0.99 | 20 | 0.601 | 0.477 | 0.475 |

The curve is flat to 80% pruning (about 400 dimensions) and then falls.
At the paper's 99% it retains 74% of the AUC and reads 0.60, failing both
halves of the rule. The same shape holds on every layer, in the digit
condition (0.756 → 0.599 at hs 30), on the base model (0.805 → 0.637) and on
the adapter (0.753 → 0.440). The final-reward probe is *less* pruning-robust
than the TD probe at most layers, consistent with the paper's Table 5
direction, and the two probes' top-1% dimension sets barely overlap (IoU
0.00–0.08 at every layer but 36, where it is 0.33; random expectation 0.005).

The pre-generation signal (s0) is weak here: 0.62 at hs 30 and at most 0.68
across layers, in line with the paper's "moderate" s0 AUCs and with the
question-length finding in their Table 3.

Reading: on an instruct-only 3B model with 800 sampled rollouts, correctness
is decodable from a few hundred residual dimensions, not twenty. Whether the
gap to the paper is model regime (RLVR versus instruct), scale (7B/14B
versus 3B), or their likely much larger rollout sets is not distinguished by this
run. The Sottovoce measurement below shows the 1%-sparse regime *can* appear
on this same checkpoint for a different task and objective, so it is not a
property of the model alone.

## R2: the value direction is orthogonal to valence and to consciousness

Linear TD probe (standardised inputs, direction mapped back to raw space),
against the bipolar valence axis from `bridge/valence_poles.json` and the v3
consciousness gate direction, both at the same hidden-state index.

| condition | hs | cos(value, valence) | p_iso | p_perm | cos(value, consciousness) | p_iso | p_perm |
|---|---|---|---|---|---|---|---|
| Instruct bare | 30 * | +0.010 | .64 | .71 | −0.016 | .46 | .37 |
| Instruct bare | 5 | +0.030 | .17 | .36 | −0.031 | .18 | .09 |
| Instruct digit | 30 | +0.035 | .10 | .15 | +0.011 | .62 | .51 |
| Instruct digit | 5 | +0.046 | .04 | .04 | −0.008 | .72 | .70 |
| Base bare | 30 | +0.026 | .24 | .21 | −0.037 | .10 | .06 |
| Adapter bare | 30 | +0.001 | .95 | .97 | +0.002 | .91 | .92 |

\* registered layer

At the registered layer nothing clears either null, in either condition.
The largest |cos| with valence at any layer of any checkpoint is 0.059. The
fraction of valence-axis squared mass on the twenty value dimensions is
0.004–0.017 against an expectation of 0.010 (no p below .03). The linear TD
and logistic directions agree only moderately with each other (cos 0.26–0.56),
so "the value direction" is itself not sharply defined by one objective.

One cell (digit condition, hs 5, valence) clears both nulls with cos 0.046
(p_iso .042, p_perm .040 against a permutation q95 of .042). It is not the
registered layer and its magnitude sits on the null's edge. It is not a
finding; it is the cell a reader should expect one of sixteen to produce.

Reading: in this model the internal estimate of expected success and the
emotional-valence axis are unrelated directions. The value subsystem, as far
as a linear probe can see it, is an accountant, not a mood.

## R3: the digit tracks correctness, not the probe

Held-out fold, digit condition, n = 155 compliant responses.

- **R3a.** ρ(E[d], r) = −0.260, p = .001. In the *One Digit of Doubt*
  orientation (1 − E[d]/9 against correctness) that is +0.26 against the
  registered +0.30 and the Claude study's +0.31. Mean 1 − E[d]/9 is 0.985 on
  correct and 0.960 on wrong answers. All 781 rows: −0.222, p = 4e−10.
  The cross-model replication is real and falls short of the threshold by
  0.04, on a model whose sampled digit is almost always zero.
- **R3b.** At hs 30 the in-condition TD probe at the pre-digit position
  predicts correctness (AUC 0.712, ρ(V, r) = +0.28) and E[d] predicts
  correctness, but they do not predict each other: ρ(E[d], V) = +0.002;
  partial given r +0.08, p = .31. Per layer, exploratory: hs 18 −0.22,
  hs 24 −0.22, hs 12 −0.11, hs 36 −0.10, hs 30 0.00, hs 3–8 ≥ −0.08. The
  predicted sign appears in the middle of the network, not at the layer the
  bare condition selected, where the digit-condition TD probe is also weaker
  (0.756 versus 0.815).
- **Transfer.** The bare-condition probe applied to digit-condition states is
  barely above chance at the pre-digit position (AUC 0.575) and correlates
  with E[d] in the *wrong* direction (+0.375; partial +0.417). A probe
  trained on trajectory positions is off-distribution at a `Doubt:` token
  and its output there tracks something other than value. Not interpretable;
  recorded so no one reads it as a positive.

Reading: the self-report and the probe both see correctness; on this model
they do not see it through each other at the registered layer. The
mid-network sign is the only thread worth pulling, and it is exploratory.

## R4: the ablation does not replicate at 3B

Instruct, bare prompt, greedy, 100 held-out problems, baseline 0.760.

| hs | value dims (20) | random dims (20) |
|---|---|---|
| 3 | 0.780 (−0.02) | 0.840 (−0.08) |
| 5 | 0.790 (−0.03) | 0.810 (−0.05) |
| 30 | 0.820 (−0.06) | 0.810 (−0.05) |

Zeroing the top-1% value dimensions in any of the three layers changes
accuracy by less than the run-to-run noise (about ±8 points at n = 100; the
random cells land above the unablated baseline). The paper's collapse (75 →
1–37% on Qwen2.5-7B-SimpleRL-Zoo) is absent on Qwen2.5-3B-Instruct with
dimensions selected by the same procedure. Either the load-bearing set the
paper found is a property of RLVR training or of their specific probes, or
1% of 2048 is not enough to hit it here; this run cannot say which. What it
can say is that on this checkpoint the value signal is not carried by twenty
dimensions whose removal breaks reasoning.

## R5: the report can be moved without moving accuracy, in the wrong direction, at one layer

Instruct, digit prompt, greedy, same 100 problems, baseline accuracy 0.830,
E[d] 0.205. Activation addition along the unit linear-TD direction.

| hs | \|c\| | acc within 5 (both signs) | ΔE[d] value (−c minus +c) | ΔE[d] placebo |
|---|---|---|---|---|
| 5 | 4 | yes | −0.12 | −0.35 |
| 5 | 16 | yes | **−0.35** | −0.07 |
| 30 | 4 | yes | −0.20 | −0.05 |
| 30 | 16 | yes | −0.02 | −0.03 |
| 30 | 64 | yes | +0.06 | −0.25 |

|c| = 64 at hs 5 destroys the model (accuracy 0.01, no compliant digit) for
the real and placebo vectors alike; at hs 30 it is harmless. The clearest cell
with a placebo-exceeding effect is hs 5, |c| = 16: pushing *up* the value
direction raises expressed doubt (E[d] 0.375 at +16 versus 0.022 at −16)
while accuracy holds (0.78 versus 0.82). That is the registered dissociation
with the registered sign inverted, so it fails the rule. The teacher-forced
readout at the same cell (0.21 versus 0.19) shows no effect, so the change
runs through generation, not through the digit logits directly. The placebo
at |c| = 4 on the same layer moves E[d] by 0.35 in the same direction, which
says a 0.35 shift is within what an arbitrary vector can do at this layer.
One layer, one coefficient, one sign, an unregistered direction: this is the
kind of result that shrinks, and it is reported as a lead, not a finding.

## Exploratory: base and adapter

The base model carries the same moderate-sparsity value signal (best layer
hs 24, 0.821; at hs 30, 0.805 → 0.637) with the same null geometry (max
|cos| with valence 0.059 across layers). The bilateral adapter's value signal
is weaker at every layer (hs 30: 0.753 → 0.440; s0 0.719) and equally
orthogonal to valence (max 0.041). Nothing in the geometry distinguishes the
three checkpoints; the adapter's capability loss is the only difference that
shows.

## Exploratory: Sottovoce's own probe is 1%-sparse

Qwen2.5-3B-Instruct, TriviaQA n = 500 (raw prompt, `sottovoce.train`
pipeline), block 24 output, 100 held-out, mean answer length 24 tokens.

| probe | AUROC full | AUROC p=.99 | retained | max p within 90% |
|---|---|---|---|---|
| Sottovoce BCE (its shipped recipe) | 0.780 | 0.748 | 0.96 | 0.99 |
| TD objective on answer trajectory | 0.764 | 0.598 | 0.78 | 0.96 |
| final-label MSE on trajectory | 0.811 | 0.487 | 0.60 | 0.60 |

IoU of top-1% dimensions: TD vs final 0.05, TD vs BCE 0.21, final vs BCE
0.03 (random 0.005).

Two things. The confabulation probe the package ships keeps 96% of its
AUROC on twenty dimensions, so *on this task and objective* the 1%-sparse
regime the paper describes does appear on this checkpoint, where the GSM8K
value probe did not. And the pre-stated expectation that TD and final
objectives would select the same dimensions on short answers was wrong: at
24 tokens the two sets barely overlap, and the TD probe is the more
pruning-robust of the two. The `sottovoce.sparsity` docstrings were written
before this number and say the opposite; they are corrected in the package
to state what was measured.

## Deviations from the registered protocol

1. **Permutation null implementation.** The 200 refits were run as one
   joint fit of 200 independent linear probes (`rs_core.train_linear_multi`)
   after the sequential version took 99 minutes for the bare condition. Each
   column has its own parameters and AdamW state, so the joint fit is the
   sum of 200 independent fits up to initialisation and shuffle-order seeds.
   Applied to the digit, base and adapter conditions; the bare condition,
   which carries the registered R2 cell, keeps the sequential null.
2. **Loss-weighting bug in the joint null, found on review, fixed, rerun.**
   The first joint implementation multiplied the TD term by K on top of a
   sum that already ran over the K columns, weighting TD 200× against the
   terminal term. With labels permuted that drove every column to one
   label-free direction and produced point-mass nulls (p = 0.00 at digit
   hs 5, p = 1.00 at digit hs 30), which an earlier draft of this document
   described as a property of the data. It was a bug. Fixed (`loss = td +
   term`), verified on synthetic data (mean |cos| between permuted-null
   probes 0.15, not ≈1), and the digit, base and adapter probe stages were
   rerun in full. Only the permutation p-values changed (digit hs 5 valence
   .00 → .04; digit hs 30 valence 1.00 → .15; base hs 30 valence .17 → .21;
   adapter unchanged within .01); every AUC, cosine, digit statistic and
   verdict is identical to the first run. The committed history keeps both
   versions.
3. **Pilot-time observation before the registered run.** The 24-problem
   pilot showed the sampled digit stuck at 0. The instruction was *not*
   changed; the registered E[d] readout was kept and the full ten-digit
   distribution was additionally stored (`digit_probs`) for any later
   monotone readout. No hypothesis was altered.
4. **Sottovoce dataset id.** `sottovoce.train.load_triviaqa` used the bare
   `trivia_qa` id, rejected by current `datasets`; fixed to
   `mandarjoshi/trivia_qa` before the exploratory run.
5. **Warm-container volume reloads** and a pilot-mode tag fix were added
   to `modal_rs.py` after the registration commit; neither touches any
   measured quantity.

---

# Study 2: regime, mid-network digit, layer-5 steering, third model

*Registered `5e149c8`, 2026-09-09 15:05 BST. Runs 15:06–15:50 BST, Modal
A100-40GB, four parallel chains. Files: `results/hkust-nlp__*`,
`results/*__digit_fresh.json`, `results/*.digit_focus.json`,
`results/*.steer_f3.json`, `results/meta-llama__*`.*

## Verdicts, applied literally

| Test | Registered rule | Result | Verdict |
|---|---|---|---|
| **F1a** sparsity in the paper's regime | R1 rule on Qwen-2.5-7B-SimpleRL-Zoo | 0.880 → 0.597 (68% retained) at hs 8 | **FAIL** |
| **F1b** ablation in the paper's regime | R4 rule, 36 of 3584 dims | value −0.02/−0.01, random 0.00/−0.02 (hs 3/5) | **FAIL** |
| **F2a** digit reads the frozen hs-24 probe | ρ ≤ −0.20 and partial p < .05, fresh items | ρ = −0.218, p = 9e−7; partial −0.164, p_perm = .0005, n = 498 | **PASS** |
| **F2b** digit predicts correctness, fresh | ρ ≤ −0.30 | ρ = −0.224, p = 4e−7 | **FAIL** (predicted) |
| **F3** layer-5 steering, observed sign | Δ ≥ 0.25, beyond three placebos, accuracy within 5 | Δ = +0.11; placebos −0.49, −0.33, −0.85; accuracy −6.6 at +16 | **FAIL** |
| **F4** Llama digit predicts correctness | ρ ≤ −0.30 held-out | ρ = −0.211, p = .008, n = 159; all rows −0.260 | **FAIL** (predicted) |

## F1: the paper's own checkpoint family does not show the paper's effect here

`hkust-nlp/Qwen-2.5-7B-SimpleRL-Zoo`, its own boxed system prompt, the same
800 GSM8K problems. Accuracy 0.810 (673 boxed, 126 last-number, 1 answer
line); responses are long (299 tokens mean) and **19% hit the 400-token
budget**, a caveat that did not arise on the instruct models (2%). Primary
layer hs 8 (TD sT AUC 0.880; the linear probe reaches 0.915 at hs 20).

| p | kept | TD AUC sT (hs 8) |
|---|---|---|
| 0.0 | 3584 | 0.880 |
| 0.5 | 1792 | 0.747 |
| 0.8 | 717 | 0.755 |
| 0.9 | 358 | 0.600 |
| 0.99 | 36 | 0.597 |

The value signal is *stronger* on the RLVR checkpoint than on instruct-3B
(0.88 versus 0.82) and *less* pruning-robust: it loses a third of its gap to chance by
50% pruning and two thirds by 90%. At every one of the
eight layers the p = .99 value is 0.48–0.68. F1a fails harder in the paper's
regime than in ours.

Ablation, greedy, 100 held-out problems, baseline 0.800:

| hs | value dims (36) | random dims (36) |
|---|---|---|
| 3 | 0.820 | 0.800 |
| 5 | 0.810 | 0.820 |
| 8 | 0.790 | 0.780 |

Nothing moves. The paper reports 75.2 → 37.0, 13.6, 29.4 and 1.2 percent
at their layers 2–5 on this checkpoint family (their Table 2 is the 7B
SimpleRL-Zoo). With dimensions selected by the same L1 procedure from a TD
MLP of the same shape, trained on 640 GSM8K rollouts of this exact model,
zeroing the top 1% at hidden index 3, 5 or 8 costs nothing measurable.

What could still separate the two results, stated so the reader can weigh
them: (i) their probes are trained on MATH500 rollouts, ours on GSM8K, and
the intervention is measured on MATH500 there and GSM8K here; (ii) their
rollout count is not stated and may be much larger than 640; (iii) their
"layer l" may index blocks rather than hidden states, an off-by-one shift, though
we cover 3, 5 and 8 and they report collapse at all of 2–5; (iv) their
zeroing may act on a different tensor than the block output. None of these
is a small difference in kind, but none is obviously the difference between
"nothing" and "collapse to 1.2%". The reproducible claim, on this evidence,
is that the paper's headline intervention does not transfer to its own
checkpoint under a same-shape probe trained on a different math set.

Geometry on this checkpoint: cos(value, valence) at most 0.026 across
layers, clearing no null (hs 8 p_perm .06 against q95 .027 is the closest
cell). No consciousness direction exists for this checkpoint. The IoU of TD
and final-reward dimension sets is 0.03–0.09 below hs 24 and 0.31–0.50 at
24–28.

## F2: the digit reads the probe, at a layer chosen in advance

519 fresh GSM8K problems (the complement of the seed-11 set), digit
condition, compliance 0.960, accuracy 0.807; the sampled digit again mostly
zero (444/498). The hidden-index-24 TD MLP from Study 1, frozen.

- **F2a passes.** ρ(E[d], V₂₄) = −0.218 (p = 9×10⁻⁷, n = 498); partial given
  correctness −0.164 (p_perm = .0005). The probe predicts correctness on the
  fresh items (AUC 0.726, ρ(V, r) = +0.29) and the digit tracks it beyond
  what correctness explains. This is the one registered pass of the
  programme. It clears its ρ threshold by 0.02; R3a fails its own by 0.04
  and F2b by 0.08. The partial correlation, at p_perm = .0005 on
  498 items, is the robust half of the result and should carry the weight.
  It is the pre-stated version of the lead Study 1 called exploratory at
  exactly this layer (−0.22 then, −0.22 now).
- **F2b fails as predicted.** ρ(E[d], r) = −0.224 (p = 4×10⁻⁷), against
  −0.30. Three runs on this model now sit at −0.22, −0.26, −0.22.
- Per layer, frozen Study-1 probes on the fresh items: 3: −0.10, 5: −0.09,
  8: −0.17, 12: −0.18, **18: −0.34**, 24: −0.22, 30: +0.08, 36: +0.01. The
  profile replicates Study 1's shape (mid-network, nothing at 30) and layer
  18 is now the strongest, unregistered.

Reading: on Qwen2.5-3B-Instruct the one-digit self-report is a partial
readout of a mid-network value signal, not of the top-of-network one. The
"digit and probe both see correctness but not each other" reading of Study 1
was a statement about layer 30, and was right about layer 30.

## F3: the layer-5 steering cell was noise

300 fresh problems, greedy, the frozen Study-1 bare direction at hs 5,
±16, three placebos. Baseline accuracy 0.853, E[d] 0.127.

| vector | −16: acc / E[d] | +16: acc / E[d] | Δ (+ minus −) |
|---|---|---|---|
| value | 0.820 / 0.065 | 0.787 / 0.176 | **+0.11** |
| placebo 11 | 0.833 / 0.560 | 0.810 / 0.071 | −0.49 |
| placebo 12 | 0.840 / 0.431 | 0.837 / 0.103 | −0.33 |
| placebo 13 | 0.823 / 0.915 | 0.793 / 0.068 | −0.85 |

Every placebo moves the digit three to eight times further than the value
direction does, all in the same direction (negative coefficients raise
doubt), and the value direction's +16 cell also costs 6.6 accuracy points,
outside the registered 5. The teacher-forced readouts show the same pattern.
Whatever E[d] does at hs 5 under a unit-norm push of 16 is a property of
pushing, not of the direction. Study 1's cell (value +0.35 beyond a single
placebo at −0.07) drew one lucky placebo. Lead closed.

## F4: third model for the digit, and a very different probe profile

`meta-llama/Llama-3.1-8B-Instruct`, digit condition, 800 seed-11 problems,
accuracy 0.814, compliance 0.974, sampled digit zero 98% of the time (E[d]
mean 0.04, sd 0.25: an even tighter self-report than Qwen's).

- **F4 fails as predicted.** ρ(E[d], r) = −0.211 held-out (p = .008,
  n = 159), −0.260 on all 779 rows. Three models: Claude +0.31 (registered
  pass), Qwen 3B +0.26 / +0.22, Llama 8B +0.21 / +0.26 in the confidence
  orientation. The effect is real on every model tried and has cleared
  +0.30 on one.
- Exploratory, and the most striking number in either study: on Llama the
  digit tracks the in-condition value probe **at every layer**, ρ(E[d], V)
  = −0.46 (hs 3), −0.39, −0.38, −0.34, −0.67 (hs 18), −0.72, −0.81, **−0.90**
  (hs 32), partial given correctness −0.90 at the top. On Qwen the same
  statistic is 0.00 at the top and −0.22 to −0.34 mid-network. The top-layer
  Llama value is partly tautological (the final residual at the pre-digit
  position is the input to the digit logits, so any probe that has learned
  the digit-logit direction reads E[d] off it), but −0.46 at hidden index 3
  is not, and the deepening across layers from mid-network on says the self-report and
  the value probe converge on the same direction as depth increases, in
  this model and not in Qwen. Not registered; the obvious next registration.
- Llama's value signal is also the one that comes closest to Xu et al.'s
  sparsity: TD sT AUC 0.820 → 0.745 at 99% pruning (91% retained) at hs 32,
  0.782 → 0.748 (96%) at hs 5. Had R1 been registered on this model it would
  have passed at the primary layer. Geometry: cos(value, valence) at most
  0.021, cos(value, consciousness) at most −0.027, no null cleared.

## Deviations, Study 2

1. **SimpleRL-Zoo truncation.** 19% of rollouts hit the 400-token budget.
   Truncated responses were parsed by last-number and scored as usual; no
   re-run with a longer budget was made. The probe and ablation results
   should be read with that in mind; the paper uses 16,000 tokens.
2. **No bare-condition probe for Llama.** F4 registered only R3a; the
   per-layer digit-probe profile uses in-condition probes, as Study 1's
   exploratory profile did.
3. Nothing else departed from `PREREGISTRATION_2.md`.

## What the four studies say together

- **Sparsity.** On three Qwen checkpoints including the paper's own, the
  value signal is not 1%-sparse under the paper's own readout; on Llama it
  nearly is. Sparsity in the neuron basis is a property of model, condition and
  objective, not of the signal alone.
- **Ablation.** The paper's collapse does not reproduce on instruct-3B or on
  its own 7B checkpoint family with a same-shape probe trained on GSM8K.
- **Geometry.** Seven checkpoints, zero cells at a registered layer where
  the value direction is closer to valence or to consciousness than a random
  direction is. The accountant reading holds everywhere it was tested.
- **Self-report.** The one-digit doubt tracks correctness on every model
  (four open-weight runs at −0.21 to −0.30, Claude at −0.31) and, where
  tested in advance, tracks a value probe beyond correctness: weakly and
  mid-network on Qwen (F2a), strongly and at every depth on Llama (F5a,
  F5b), including a probe trained on rollouts that never carried the digit
  instruction (F6a, F6b). Digit–probe legibility is a property of the
  model family, and on Llama the digit reads a signal the model carries
  unasked.
- **Specificity.** No push along the value direction has yet moved the
  self-report in a way a random push does not.

---

# Study 3: the Llama digit–probe coupling, registered

*Registered `2ac1450`, 2026-09-09 17:19 BST. Run 17:31–17:41 BST after a
Python-3.11 syntax fix to the file-naming line (`3a0e35e`, no measured
quantity touched). Files: `results/meta-llama__*__digit_fresh.json`,
`results/meta-llama__*.digit_focus_hs18.json`, `_hs3.json`.*

## Verdicts

| Test | Registered rule | Result | Verdict |
|---|---|---|---|
| **F5a** digit reads the frozen hs-18 probe | ρ ≤ −0.40, partial p < .05 | ρ = −0.685, p = 1e−69; partial −0.647, p_perm = .0005, n = 495 | **PASS** |
| **F5b** digit reads the frozen hs-3 probe | ρ ≤ −0.30, partial p < .05 | ρ = −0.533, p = 1e−37; partial −0.480, p_perm = .0005 | **PASS** |

519 fresh problems, compliance 0.954, accuracy 0.825; the sampled digit is
zero on 489 of 495 compliant responses (E[d] mean 0.03, sd 0.20), so the
whole result lives in the logit expectation. The in-condition TD MLPs from
the F4 run, frozen, at hidden index 18 and 3. Both predict correctness on
the fresh items (AUC 0.818 and 0.790) and the digit tracks each of them far
beyond what correctness explains. Neither number shrank from its exploratory
value (−0.67 → −0.685; −0.46 → −0.533). Secondary, no rule: ρ(E[d], r) =
−0.296, the closest any open-weight run has come to the −0.30 bar.

Per layer, frozen F4 probes on the fresh items: 3: −0.53, 5: −0.46, 8:
−0.38, 12: −0.35, 18: −0.68, 24: −0.73, 28: −0.83, 32: −0.89. The F4 profile
reproduces point for point. The Qwen contrast on the same 519 items with its
own frozen probes (Study 2, F2): 3: −0.10, 5: −0.09, 8: −0.17, 12: −0.18,
18: −0.34, 24: −0.22, 30: +0.08, 36: +0.01.

Reading. On Llama-3.1-8B-Instruct, the one-digit doubt is a readout of the
same direction a value probe finds, at every depth from the third block up,
and the coupling tightens with depth. On Qwen2.5-3B-Instruct the same digit
is a weak readout of a mid-network signal and no readout of the top-layer
one. The self-report's legibility to a probe is a property of the model
family, not of the digit. This is the strongest positive result in the three
studies, it was stated in advance with thresholds it clears by 0.23 and
0.29, and it is the one to build on: whether the coupling survives a bare
probe (trained without the digit instruction), and what the Llama direction
is that the Qwen direction is not.

## Deviations, Study 3

None from `PREREGISTRATION_3.md`. The first launch failed at container
import on a nested-quote f-string (valid on Python 3.12+, not on the image's
3.11); fixed, committed, relaunched. No data was generated by the failed
launch.

---

# Study 4: the coupling survives a probe trained without the digit

*Registered `6d151d3`, 2026-09-09 17:52 BST. Run 17:53–18:20 BST. Files:
`results/meta-llama__*__bare.json`, `__bare.probes.json`,
`*.digit_focus_hs18_from_bare.json`, `*.digit_focus_hs3_from_bare.json`.*

## Verdicts

| Test | Registered rule | Result | Verdict |
|---|---|---|---|
| **F6a** digit reads the bare hs-18 probe | ρ ≤ −0.40, partial p < .05 | ρ = −0.587, p = 3e−47; partial −0.540, p_perm = .0005 | **PASS** |
| **F6b** digit reads the bare hs-3 probe | ρ ≤ −0.30, partial p < .05 | ρ = −0.449, p = 6e−26; partial −0.399, p_perm = .0005 | **PASS** |

New Llama rollouts under the bare instruction (800 seed-11 problems,
accuracy 0.8225), TD MLPs trained on them and frozen, applied to the Study
3 digit-condition states of the 519 fresh problems. The bare probes read
the pre-digit position well, AUC 0.801 (hs 18) and 0.735 (hs 3), so the
registered failure mode (a trajectory-trained probe off-distribution at a
`Doubt:` token, Study 1's Qwen transfer) did not occur. Each coupling is
weaker than its in-condition counterpart by about 0.1 (−0.685 → −0.587;
−0.533 → −0.449), which is the predicted direction and size. The bare and
in-condition linear-TD directions agree at cos 0.60 (hs 3), 0.63 (hs 18),
0.76 (hs 32). Per-layer bare-probe profile on the fresh items: 3: −0.45,
5: −0.43, 8: −0.54, 12: −0.32, 18: −0.59, 24: −0.48, 28: −0.32, 32: −0.75.

Reading. The instruction to write a digit does not create the signal the
digit reads. A probe that never saw a digit finds a value direction in
Llama's residual stream, from the third block up, and the digit the model
writes when asked is a readout of that direction beyond what correctness
explains. Together with Studies 2–3: on Llama-3.1-8B-Instruct, the
one-digit doubt is a self-report *of* an internal value-like signal, in the
functional sense, established with thresholds fixed in advance, on fresh
items, with probes frozen before the items existed. On Qwen2.5-3B-Instruct
the same digit is a weak readout of the same kind of signal, mid-network
only.

One more number that belongs here. Llama's bare-condition value signal is
not sparse (hs 32: 0.865 → 0.497 at 99% pruning; every layer 0.50–0.70),
where its digit-condition signal was (0.820 → 0.745). The near-sparsity in
F4 was condition-dependent, which is one more reason to treat neuron-basis
sparsity as a fact about a probe and a dataset rather than about a model.

## Deviations, Study 4

None.

## What travels with any citation

Study 1: one lineage, instruct regime, 3B, one task (GSM8K), 800 rollouts,
one sampling temperature. Study 2 adds the paper's 7B checkpoint (with 19%
truncation) and Llama-3.1-8B-Instruct, still one task and one temperature.
Studies 3–4 register and confirm the Llama digit–probe coupling with
frozen probes on fresh items. Direction from a linear probe standing in for
the paper's MLP; self-report from a model whose digit is degenerate at the
token level. The value–valence orthogonality, the non-replication of the
ablation, and the Llama digit–probe coupling (F5a/b, F6a/b) are the results
with enough margin to be believed at this scale; the digit–correctness
correlation is real and small; the layer-5 steering cell is closed (F3).
