Quasiqualia
Research note · Preliminary

Four Results the Harness Made Up

Four findings from our own interpretability work on Qwen2.5-7B-Instruct came from how the experiments were run and analyzed, not from the model. In three, the comparison rested on identical copies (in one, 20 of 20 “treated” replies matched their controls exactly); the fourth measured its own intervention. All four are now retracted.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

This is a self-correction, not a new experiment. Four studies on one open-weight model (Qwen/Qwen2.5-7B-Instruct) ran by May 2026 (results committed 2026-05-11) and were audited and retracted in August 2026; none was ever released publicly, and the result numbers below were recomputed from the stored per-trial and result files. No pre-registration record exists for any of the four, and no LLM judge was used anywhere (one study counted words with a fixed string list). The biggest caveat: the questions these studies asked are now open, not answered in the other direction.

An intervention experiment has two arms: one where you do something to the model, and one where you don’t. The finding is the difference. If the harness (the code that runs and analyzes the experiment) fails to do the thing, the arms come out identical, which looks exactly like a clean null result.

We recorded four results as findings in the internal records of a programme that reads and steers directions inside an open-weight model. An audit in August 2026 found that all four were made by the harness. Three rested on comparisons of numbers or text that were, wholly or almost wholly, the same. The fourth reported an effect that the steering vector would produce by simple arithmetic, with a p-value that cannot exist at its sample size. The lesson applies to anyone running activation interventions, and the cheapest fix is one line.

01 · Setup

What the four studies did

All four used Qwen/Qwen2.5-7B-Instruct, steering or probing at layer 24 of its 28 layers (the spillover study read its directions further downstream, at layer 27). The programme had extracted labeled directions in the model’s activations: self-monitoring, groundedness, engagement, and three for involvement, creativity and multi-perspective processing. A “projection” is how far an activation points along one of those directions. Decoding was greedy except in the drift study, which was a single conversation sampled at temperature 0.7 (top-p 0.9) with no fixed seed.

  • Knockout study. A system line of labeled state values (“[Your state: R:+3.2 P:-1.5 Q:+4.1 G:-2.3 V:+1.8]”; the letters are the programme’s shorthand for its directions) was placed before 20 questions about the model’s inner state. The knockout arm was meant to stop the model attending to those labels. Self-reference was scored by counting which of nine strings (“i “, “my “, “feel”, “aware” and so on) appeared in a reply of up to 200 tokens.
  • Recovery study. The model was pushed along the self-monitoring direction for the first 10 of 200 generated tokens, at seven strengths from -3 to +3 on 5 prompts, then left alone for 190 tokens.
  • Drift study. A 200-turn conversation, with groundedness and engagement projections read at turns 1, 5, 10, 20, 50, 100, 150 and 200.
  • Spillover study. The model was steered along the engagement direction at five strengths from -2 to +2 while generating up to 100 tokens for 20 neutral prompts, and the three other directions were read at layer 27, three layers later (the script asked for four, but the model has only 28 layers, so it silently used the last one).
What it found
  • Knockout: all 20 knockout replies are byte-identical to their normal counterparts (SHA-256 match), and every one of the 40 trial files records 0 label positions found. The knockout never ran.
  • Recovery: in 28 of the 30 steered cells, the 190-token recovery trace is bit-identical to the unsteered control.
  • Drift: the measurements at turns 50, 100, 150 and 200 are one number recorded four times. On the 5 distinct points, the groundedness correlation with turn is -0.10 and the engagement correlation is 0.00.
  • Spillover: the observed spillover slopes are 0.65, 0.84 and 0.64 of what plain arithmetic leakage predicts, and the reported p-value is below the smallest value the test can produce.
02 · Failures

Four ways to get a result from nothing

The knockout that never fired. The code searched each token’s text for strings like “R:” or “+3”. The tokenizer splits the state line into pieces that contain none of them, so the list of positions to knock out was always empty. The hook (a small function attached to a layer to change its output) that would have done the knockout is registered only if that list is non-empty, so it was never attached. Two more defects would each have voided it anyway: the generation call returned no attention weights to edit, and editing weights after the attention output is computed cannot change that output. The knockout arm ran the normal arm’s code path.

The steering that ended before the measurement. Nothing checked that the push changed anything that lasted. During the 10 steered tokens the projection did differ from control in all 30 cells, so the steering ran. But it could reach the recovery window only through which tokens the model chose, and in 28 of 30 cells it must have chosen the same ones. The aggregate shows it: mean recovery is -19.363 at every strength from -3 to +1. The recorded test for a second stable state compared that number with an identical copy of itself, and the difference was exactly 0.

The conversation the probe stopped reading. The script capped the input at 3,800 tokens and silently cut off everything after that. By turn 50 the conversation (23,920 characters) was past the cap, so every later probe read the same opening and returned the same activation: groundedness -22.155 and engagement +49.419 at turns 50, 100, 150 and 200. The input the model replied to was cut the same way: the stored reply openings at those four turns are word-for-word the same (“multi-faceted approach is necessary”), which suggests the model never saw the later questions. Those four copies form the monotone tail that made a trend. The script also silently skipped any direction whose vectors it could not load, so only 2 of the 5 it planned (groundedness and engagement) were measured. The result had been flagged as marginal a month before the audit; corrected for testing two directions, the exact p is 0.124. A recommendation to monitor groundedness in long sessions, which rested on this trend, was withdrawn.

The spillover that was arithmetic. A steering vector added to the activations stays in everything downstream unless something removes it. The script’s comments named this risk and tried to avoid it by reading further downstream (three layers, not the four it requested), but the later activation still contains the added vector. Because the engagement direction has cosines of 0.607, 0.542 and 0.537 with the three other directions, pure pass-through predicts their projections should rise by about that much per unit of steering. The observed slopes, averaged over the 20 prompts, were 0.392, 0.455 and 0.345. Reading later removed some leakage (about 16 to 36% of it), not all; nothing in the observed shift requires a separate causal effect, and nothing here rules one out. The recorded statistic was a rank correlation of 1.000 on five averages, with p = 1.4e-24. With five points there are 120 possible orderings, so the smallest p an exact test can give is 2/120 = 0.017 two-sided (1/120 = 0.0083 one-sided). The tiny value is an approximation that fails at a perfect correlation.

03 · Retractions

What was recorded and what the files show

Study Recorded claim (retracted) What the files show Direction of the error
Knockout Knockout has no effect on self-reference, d = 0.000, p = 0.506 20 of 20 reply pairs identical; that p was one-sided, and the two-sided p on identical data is 1.0 Manufactured a null
Recovery One stable state; recovery identical at all strengths 28 of 30 recovery traces are the control Manufactured a null
Drift Groundedness falls over 200 turns, r = -0.710, p = 0.048 4 of 8 points are one measurement; exact p 0.062 even before that Manufactured a trend from a frozen probe
Spillover Engagement is a causal hub, r = 1.000, p = 1.4e-24 Shift fits leakage; p is below the 0.017 exact floor (two-sided) Manufactured an effect

One correction to our own summary: it described the engagement trend in the drift study as r = +0.710, p = 0.048. The stored file says r = +0.685, p = 0.061, and the script itself had marked it as no drift.

The spillover study does contain a small, real shift. Comparing +2 with -2 steering per prompt (n = 20), the three projections rose by 1.53 [95% bootstrap interval 0.96, 2.04], 1.78 [1.55, 2.06] and 1.38 [0.95, 1.79], standardized effects (d_z) of 1.21, 2.96 and 1.42, about 4 to 10% of their baseline values. That shift is within what leakage alone predicts: the fitted slopes are 0.64 to 0.84 of the predicted ones. It is not evidence of a hub.

04 · Lesson

The check that would have caught three of them

Three of the four failures compared things that did not differ. A broken intervention behind a null is the dangerous kind: a null is easy to believe, and an intervention that never happened says nothing about whether the thing matters. A one-line assertion that the treated output (tokens, text, or the probe’s input) differs from the control would have stopped all three before any statistic was computed. For truncation, the equivalent is to record the length of the input the probe saw and fail if it stops changing.

The fourth needs a different control. Steering along a direction that overlaps the readout moves the readout whatever the model does. The test that separates leakage from a causal effect is a random direction with the same cosine to the readout: leakage predicts the same slope for it, a genuine hub predicts roughly none.

05 · Limits

What this does not show

These retractions do not show the opposite claims. Whether the labels drive self-reference, whether the self-monitoring direction has one stable state or two, whether anything drifts across long conversations, and whether engagement drives the other directions are now unmeasured. The spillover study’s matched random-direction arm was never run, and its cosines came from local copies of the direction files (they match the audit’s values). All trial files were present and none were dropped. A re-run would need a verified knockout target list, steering that lasts through the measurement window, a logged context cap, exact p-values at small n, and a manipulation check in every arm.

Data and code

Where the evidence lives

Experiments AY71 (the knockout study), AY68 (the recovery study), AY73 (the drift study) and AY64 (the spillover study). Scripts: research/experiments/modal_ay71_attention_knockout.py, modal_ay68_r_stability.py, modal_ay73_longitudinal_drift.py, modal_ay64_q_steering.py. Results: research/results/ay71/ay71/ (40 trial files), research/results/ay68/ay68/ (35), research/results/ay73/ay73/aggregate.json (8 measurements), research/results/ay64/ay64/trials/ (100). Audit reanalyses: _audit/reanalysis/analyze_ay71_attention_knockout.py, analyze_ay68_r_monostable.py, analyze_ay64_q_causal_hub.py, and _audit/22_ay73_refutation.md. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026harnessmade,
  title={Four Results the Harness Made Up},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/harness-made-up.html}}
}