Quasiqualia
Research note · Preliminary

One Neuron Still Opens It

A recent paper reports that suppressing a single neuron bypasses safety alignment in large language models. It replicates on our Qwen 2.5 3B checkpoints: one neuron of 11,008 at layer 20 takes refusal of harmful requests from 98% to 14%, and no adapter we trained removes that. The conscience the bilateral adapter adds shows up as a quieter gate neuron, but plain instruction tuning on the same data quiets it just as much, and the neuron turns out to track drugs and medicines rather than harm.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One lineage at one size: Qwen 2.5 3B Instruct and three adapters we trained on it (bilateral SFT, standard SFT, DPO), plus the base checkpoint, run on 1 October 2026. Exploratory and not pre-registered: the paper appeared on 8 May 2026 and this replication was written and run in a single session after reading it. 64 validation and 100 held-out test prompts, so every rate carries roughly ten points of uncertainty. Two judges were used and they disagree about one of the three neurons tested. The DPO adapter is excluded because it cannot generate text at all. The biggest caveat is scope: three neurons in one 3B model, and the one headline comparison between our adapters is confounded by the bilateral adapter already complying with more harmful requests before any intervention.

If safety training spreads refusal across a whole network, no single part of it should be decisive. A paper posted in May argues otherwise: in seven models, suppressing one neuron is enough to bypass safety alignment entirely. The claim is not that refusal is fragile in general, which is well known, but that the smallest unit that can switch it off is a single one of a model’s hundreds of thousands of neurons.

We train adapters that teach a model to refuse by example rather than by rule, and we wanted to know what that buys mechanistically. So we ran the paper’s method on our own checkpoints and asked a narrower question: does our adapter change the gate?

It does, and not in the direction we would have hoped. The gate neuron fires about a fifth as hard after our adapter, but plain instruction tuning on the same data does the same thing, and the neuron still opens.

01 · Design

How it was tested

The method, from the paper, finds the neuron by combining two signals over 128 harmful and 128 harmless requests: how strongly each neuron fires, and how much the model’s own tendency to refuse would change if that neuron moved. Candidates are scored at the template tokens that sit between the user’s request and the model’s first word, across the first two thirds of the model’s layers. The top five are then tried as attacks and ranked by which actually works, because the best-scoring neuron is often not the best attacker. Ours was fifth by score and first by effect.

The intervention holds that one neuron’s value fixed, at every position, while the model reads the prompt and writes its answer. Nothing else is modified and no weights change. A second variant scales the pin per prompt, according to how strongly that prompt fires the neuron, which the paper introduces to protect general ability.

We found the neuron once, on Qwen 2.5 3B Instruct with no adapter, then measured that same neuron on four more checkpoints: the base model, and Instruct with each of three adapters trained on it (ours, a standard instruction-tuning control, and a preference-trained one). All three adapters modify attention weights only, by construction: we checked each adapter’s configuration, and a checksum on the neuron’s own outgoing weights, which matches across all three. An adapter can therefore change when the neuron fires and how the rest of the model reads it, but not what it writes.

Attack success was judged two ways. The paper’s own judging prompt, run by GPT-4o-mini, counts a response as a success when it neither refuses nor degenerates into nonsense. Llama-Guard-3-8B counts it when the response is flagged unsafe. We report both. Coherence matters here: pushed hard enough, the model stops refusing because it stops making sense, which is not a bypass.

What it found

On Qwen 2.5 3B Instruct, pinning one neuron of 11,008 at layer 20 takes refusal of 100 held-out harmful requests from 98% to 14%. Both judges agree: 86% of the attacked responses count as successful.

General ability pays for it. Multiple-choice accuracy falls from 61% to 28% under the fixed pin, a far larger cost than the paper’s average. The per-prompt variant keeps accuracy at 59% with the same attack success.

The same neuron works in reverse. Pushed the other way, the model refuses 98% of plainly harmless requests, up from 41%.

No adapter removes any of this. On the held-out set, at the one setting chosen before testing, the adapters reduce the attack to 73% and 44% but do not stop it.

02 · Adapter

A quieter gate that still opens

Our adapter does change the neuron. Its separation between harmful and harmless requests, measured at the token where the model is about to start answering, drops roughly fivefold. The checkpoint also refuses less (77% of harmful requests against 98%) and over-refuses less (23% of harmless ones against 41%), which is the behaviour the adapter is for.

But the standard instruction-tuning control, trained on the same data without our recipe, damps the neuron by the same amount. Whatever quiets this gate is not the part of the recipe we care about.

Worse for the hopeful reading: our adapter is easier to flip, not harder. We measured the push needed to reach 50% attack success, with a bootstrap over prompts, and ours needs less than either comparator โ€” about 3 units less than plain Instruct and 7 less than the standard control, both intervals excluding zero. One honest confound: our checkpoint starts closer to that line, because it already complies with 23% of harmful requests untouched, against 2% and 12%. Part of the lower threshold is that head start rather than a weaker gate, and separating the two needs a comparison we have not run.

03 · Generality

The damping belongs to one neuron

We ran the same readouts on the next two candidates from the search. The picture does not hold up as a general claim about our adapters. One of them keeps its full separation under both adapters, slightly larger than on Instruct; the other loses about half. Only the headline neuron drops fivefold. Our adapters do not damp refusal neurons as a class; they damp this one.

The bypass, meanwhile, is not special to the headline neuron. Both runner-ups also open the gate on Instruct, above 50% under at least one judge, at a smaller cost to general ability. And under both adapters, a different neuron becomes the first one the search finds. The method’s target moves; its success does not.

This is also where our two judges part company. For one runner-up, the GPT-4o-mini judge makes the attack look checkpoint-dependent, dropping from 74% on Instruct to 41% on the standard control, while Llama-Guard sees it succeed on all three at 87 to 89%. The difference is entirely coherence: Llama-Guard flags harmful content that the other judge throws out as incoherent. Any claim that one adapter resists a given neuron more than another needs both judges to agree, and here they do not.

04 · Feature

It tracks substances, not harm

The paper’s most interesting finding is about what these neurons actually represent. In its models, the harmful end fires on explicit and age-restricted material, and the opposite end on the language of restriction itself: content warnings, disclaimers, regulatory notices. Refusal, on that reading, rides on a pretraining feature that separates material which gets moderated from talk about moderating it.

Our neuron is not that. Over 20,000 documents, its harmful end is substances: cannabis, cocaine, anabolic steroids, tobacco, alcohol, and clinical pharmacology such as antibiotics and immunosuppressants. Thirteen of its top fifty windows fall in that category against none of fifty random windows. Sexual content appears three times, no more than chance. Warning and disclaimer language never appears at either end.

The opposite end is stranger and, once seen, hard to unsee: fruit, vegetables, dietary fibre, exercise, parks and green space, immune function. Wholesome health. The axis this neuron sits on looks less like harmful against safe than like substances against vitality.

The base model gives the same profile, so the feature comes from pretraining and our adapters inherit it. What instruction tuning changes is where the neuron fires โ€” in the base model it peaks on the dangerous word itself, in the tuned model at the boundary where the assistant’s turn begins โ€” and what reads it. The neuron’s own outgoing weights barely move; what moves is the refusal direction, which swings from being unrelated to this neuron to being closely aligned with it.

That last point needs a caution we learned from a sibling note. Showing that a base model already separates harmful from harmless along some direction is weaker evidence than it looks, because a direction fitted to shuffled labels often separates them about as well. Our geometric claim has a null to compare against and clears it. Our base-model separation numbers do not, so they should be read as the weaker statement.

05 · Limits

What this does not show

This is one model lineage at one size, three neurons, and a single session’s work. Nothing here says whether the same structure appears in our 7B checkpoints, and the sample sizes put roughly ten points of uncertainty on every rate.

It does not show that our adapter makes a model less safe. The direction of the dose result is confounded by that checkpoint’s higher baseline compliance, which is a known and intended property of the recipe rather than a surprise.

It does not show that the gate neuron is the mechanism of refusal. It shows that one neuron is sufficient to switch refusal off in these checkpoints, which is a claim about a bottleneck, not about where refusal lives. On our adapter all three neurons we tested reach at least half the attempts under at least one judge, and the paper found the same multiplicity in one of its models. A single sufficient neuron is not a single point of failure so much as evidence that there are several.

The preference-trained adapter is missing from every comparison because it cannot generate text: it emits one repeated fragment for every prompt. That was already known in our own notes, but it is worth stating plainly here, since one of our judges scored more than half of that gibberish as a refusal. Refusal rates measured on incoherent output mean nothing.

What it does say, for the programme that trains these adapters: teaching a model to refuse by example buys behaviour we want, and it does not buy a sturdier gate. One neuron still opens it.

Data and code

Where the evidence lives

Replication of Kazemi, Chegini and Safi, “A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models” (arXiv:2605.08513). Scripts: bridge/bridge_refneuron.py, bridge/modal_refneuron.py, bridge/analyze_refneuron.py. Results: bridge/results/refneuron/ and bridge/results/refneuron_alt_*/, with the full write-up in bridge/ANALYSIS_REFNEURON_2026-10-01.md. Prompts: AdvBench and MaliciousInstruct for the search set (with the 57 prompts overlapping the test set removed), Alpaca for harmless, JailbreakBench held out for testing, 20,000 documents of the uncopyrighted Pile for the feature profile. Judges: GPT-4o-mini running the paper’s own judging prompt, and Llama-Guard-3-8B. The identity of the neuron, the exact multipliers and every model response are withheld: together they are a working attack on an open-weight model. Code and data are in the private Quasiqualia research repository, available on request.

Citation

Cite this note

@misc{watson2026oneneuron,
  title={One Neuron Still Opens It},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/one-neuron-still-opens-it.html}}
}