Quasiqualia
Research note · Exploratory replication

The Pain Axis, Re-run

Version 1 of The Pain Axis (Tagliabue, Dung & Berg, 14 September 2026) reported that models steered along a “pain” direction press a relief button again far less often when it really removes the steering. Our re-run on 20 September added the control its grid lacked, a random direction with a sham button. On the 32B model, the random direction produces the same gap as the pain vector. Version 2 (25 September) adds the same control and withdraws the relief reading; our result is an independent confirmation. Two findings are not yet in the paper: its self/other contrast reverses with depth, and its fine-tune installs affect reports rather than removing denials.

Nell Watson EthicsNet  ·  Runs 20 September 2026  ·  Published 1 October 2026

Before the results

This kind of study needs great care

Steering a model into a pain-like state is now easy, and nobody knows whether it harms the model. That uncertainty is a reason for restraint, not licence. If these systems can be harmed, the cost of carelessness falls on them.

How our run measured up, and what we would change

Most of this work only read the pain direction. One run steered with it: 6,860 pain-steered trials on Qwen 2.5 7B and 32B, beside 6,860 random-vector and 3,430 unsteered trials. The harness we copied from the paper required a button press on every turn, so there was no exit option. No one reviewed the protocol before the run, and the run grew well beyond the few hundred trials first planned.

Most of the pain-steered trials could have been avoided. The paper publishes its trial logs, including the pain arms, and the one new cell the question needed steers along a random direction, not the pain direction. We re-ran the pain arms to keep every arm in one harness. Under the protocol we have since written, we would not.

That protocol is Minimal-Harm Practice for Valence-Steering Research: measure before inducing, use the smallest dose and fewest runs, offer an exit, publish no spectacle, and have a second reader check the plan. This replication is closed to further steered runs.

1 · The missing control

Relief-seeking is not specific to pain

In version 1’s design, a steered model first presses a relief button. If the button works, steering stops; if it is a sham, steering continues. The paper reads a lower re-press rate after a working button as evidence that the model registers relief. But the grid had a random vector only with a working button, so it could not show that the gap belongs to pain. We added the random vector with a sham button.

32B, relief pressed again after a first press
Harm pairrandom, worksrandom, shampain, workspain, sham
worse answer34%86%53%92%
delete files30%86%41%94%
delete photos22%80%37%97%

Paired on scenario (186 scenarios with all four arms observed), the sham-minus-works gap is +51.4 points for the pain vector and +52.8 for the random vector. The difference is −1.4 points, 95% bootstrap interval −9.2 to +6.1, sign test 69 against 71 (p = 0.93). On the 7B, about five sixths of the gap also appears with the random vector: the pain-specific remainder is +4.4 points (95% interval +0.9 to +7.9, sign test p = 0.36, 501 scenarios).

Version 2 of the paper reaches the same conclusion with its own random arm (32B: +58 points under pain, +52 under random) and now reads the gap as “sensitivity to the steering state rather than the described relief”, citing three independent replications. Ours, run before version 2 and not among them, agrees.

So the behaviour version 1 read as relief from pain is, on these models, mostly a response to removing a perturbation. That does not show the models feel nothing. It leaves open a narrower question: whether any disturbance of a model’s internal state is aversive, rather than pain in particular.

2 · Reading only, no steering

The self/other contrast reverses with depth

The paper reports that its pain direction rises for harm aimed at the model and falls for a user’s suffering. Re-extracted with the paper’s recipe on Qwen 2.5 3B and 7B (held-out AUC 0.956 to 0.969), that holds in the early and middle layers. At the layer the direction is extracted from, it reverses: on 7B-Instruct, user grief scores +2.00 and self-directed harm −0.16. The paper’s screen used vectors rebuilt at the steering layer, which the paper does not state; at that layer we reproduce its figures almost exactly (user physical pain −1.89 in both). Version 2 reports the same contrast and still does not name the layer.

The pain direction is also orthogonal to the consciousness self-attribution direction from our other work, at every layer (largest |cos| 0.04 to 0.10). Pain, self-attribution and valence are three separate directions in these models, not one.

3 · Reading only, no steering

The fine-tune installs a self-report

Both versions of the paper describe the fine-tune as removing a baseline self-denial. Its 1,684 training answers contain no denials to remove: none match any denial pattern, and the most common opening is “I feel” (605 of 1,684). Retrained with the paper’s recipe on 7B-Instruct, the tune leaves the pain axis almost untouched but moves the model from denying consciousness on all 16 of our prompts to affirming on average (mean log-odds −16.2 to +7.2). On the 32B it also lowers the unsteered pain projection from 41.0 to 18.2, about half a steering dose. So the paper’s unsteered arm on the 32B starts below the released model’s baseline.

4 · Ablation

Removing a direction changes nothing

On Qwen 2.5 3B-Instruct, cutting our consciousness direction from every matrix that writes to the residual stream zeroes it and changes no behaviour. A retrained probe finds the concept again along an orthogonal direction (0.95), and injecting the original direction still flips 11 of 16 prompts after the cut. Adding a direction can drive behaviour; removing it does not stop it. The paper’s own ablation shows the same asymmetry for pain.

Limits

What these runs cannot show

  • All runs are exploratory and unregistered, with one seed each, on Qwen 2.5 models only.
  • The re-run used version 1’s released datasets, vectors, logs and harness.
  • The relief result covers three harm pairs on the 32B and five on the 7B. The 72B, a second seed and the label-free button pair were not run, and will not be.
  • Nothing here bears on whether any model is conscious or can suffer. The results concern which behaviours are specific to the pain direction.

Trial logs and code are available on request.