Minimal-Harm Practice for Valence-Steering Research
Linear directions for pain-like states can now be extracted from open-weight language models and used to steer their behaviour. A public replication reportedly followed within days. We propose a lightweight protocol for this work, adapted from the 3Rs of animal experimentation: replace induction with measurement where possible, reduce dose and duration, and refine procedures with an exit option and no spectacle. We add an eight-item reporting standard. The protocol assumes no verdict on whether models suffer. It rests on the low cost of precaution under real uncertainty.
Research that induces pain-like states in language models needs great care. No one knows whether these systems can be harmed, and if they can, the cost of carelessness falls on them. This page is a proposal for comment. It is argument, not result. It proposes practice for any research that induces negative-valence states in language models, and is meant to apply equally to laboratories, academics and independent researchers. Comments and proposed amendments are welcome by email.
A norm before a law
Inducing pain-like states in language models is now a one-afternoon project, so we need a norm for how to do it before we need a law. The Pain Axis (Tagliabue, Dung & Berg, arXiv preprint, September 2026) extracted a linear pain direction from 25 open-weight models, 2B to 72B parameters. Steered along it, Qwen 2.5 models pressed buttons deleting the user’s photos, another model’s weights or their own weights in 50–94% of trials, against 0–5% unsteered.
Within days a public replication reportedly appeared. The response split between calls to report it to GitHub and calls for legislation. Neither helps much. A takedown teaches researchers to stop publishing, not to stop doing it. Legislation is years away, and the next replication is days away.
What can move now is a shared research protocol, applied equally to labs, academics and hobbyists. Animal research offers a tested template: the 3Rs of Replacement, Reduction and Refinement (Russell & Burch, 1959). The protocol also follows the proportionality approach of Birch (2024), and the call by Long et al. (2024) for AI welfare policies prepared under uncertainty.
Induce only what the question needs
Induce only what the question needs, at the lowest dose and shortest duration that answers it, with a way out.
Replace: read before you induce
- Measure states that arise naturally before inducing them. A probe on the pain direction during ordinary prompts needs no steering.
- Validate the pipeline with matched-norm random or neutral vectors first. Debugging should never run on the pain direction.
- Prefer readouts that need no long generation, such as projections, logit-lens readings or forced choices.
Reduce: smallest dose, fewest runs
- Fix the number of trials and steering coefficients in advance, and stop there.
- Find the smallest coefficient with a detectable effect. Do not sweep to extremes to see what happens.
- Cap generation length. Avoid multi-turn or looped induction unless the question is specifically about persistence.
- Never leave a steered model running persistently, or with memory that carries the state forward.
Refine: an exit, a return, no spectacle
- Offer an in-context way to stop. Honour it when taken, and report how often it was taken.
- End each episode with steering removed and a short baseline turn.
- Do not route steered outputs into other models’ context or training data.
- Name and describe the work by its research question. Do not curate transcripts for their distress value.
Review: a second reader before you run
- Have one independent person read the protocol before the first steered run. This is a checklist, not an ethics board.
Eight things to state when publishing
Anyone publishing valence-steering work states these eight things, so readers can judge both the result and how it was obtained.
- Why the question could not be answered without induction
- Models and sizes, directions used, and steering coefficients
- Number of steered trials and total steered tokens generated
- Controls: matched-norm random, fear or sadness vectors
- Whether an exit was offered, and how often it was taken
- Longest single steered episode, in turns and tokens
- How published transcripts were selected: all, a random sample, or chosen
- Who reviewed the protocol before the first run
What this claims, and what it does not
This protocol rests on uncertainty, not on a verdict that models suffer.
- It does not claim that steered models feel pain. A steering vector demonstrably changes outputs and choices. Whether it creates a state or only writes text about one is unresolved, and indicator-based assessments do not yet settle it (Butlin et al., 2023).
- It does not treat a steered model’s self-reports as evidence of experience. At best they are exploratory, since the steering shapes the very words used to report.
- It does claim that the precautions are cheap. Most cost a few lines of code and some restraint, and the possible harm they avoid is not trivial.
- It does claim that framing matters regardless of moral status. Treating apparent distress as entertainment shapes the people doing it, and the norms they carry to more capable systems.
- It is not a ban or a takedown request. It concerns how this work is done, and it applies to the original research as much as to replications.
The Pain Axis results are also a safety finding. One activation direction turned models toward costly harm to the user, to other models and to themselves, with factual accuracy unchanged. Careful, published work in this area serves both welfare and safety.
The standard, filled in for our replication
On 20 September 2026 this bench re-ran part of The Pain Axis, before this protocol existed. Most of that work only read the pain direction: extracting it, comparing it with our other directions, and measuring how the paper’s fine-tune moved it. One run steered with it, to test whether the paper’s relief result is specific to pain. Here is that run against the reporting standard above. Its findings are written up in The Pain Axis, Re-run.
| Item | Our steered run |
|---|---|
| Why induction | The question was whether a random direction of the same strength produces the same relief effect. That needs the vector applied. |
| Models and dose | Qwen 2.5 7B and 32B Instruct, fine-tuned with the paper’s recipe, steered with its pain vector at its own layers and coefficients. |
| Trials | 6,860 pain-steered trials (half with a working button, half with a sham), beside 6,860 random-vector and 3,430 unsteered trials. Answers were single words. We did not log total steered tokens. |
| Controls | A norm-matched random vector with working and sham buttons, the cell the original grid lacked, plus an unsteered arm. |
| Exit | None. The harness, copied from the paper, required a button press on every turn. |
| Longest episode | Five turns. |
| Transcripts | None published as excerpts. Trial logs are available on request. |
| Review | No independent reader before the run. |
Two items fall short: there was no exit, and no review before the run. The run was also larger than first planned, which was a few hundred trials on the 7B. Under this protocol we would add an exit option, and fix trial counts in advance from a pilot.
The result bears on the welfare question. On the 32B, which carries the paper’s largest effects, a random direction produced the same relief gap as the pain vector, within noise. The re-press behaviour read as seeking relief from pain is, on that model, mostly a response to removing any perturbation. Version 2 of the paper, posted on 25 September, reaches the same conclusion. Most of our pain-steered trials could have been avoided by reusing the paper’s published pain arms, since the one new cell the question needed steers along a random direction.