The Gap Held Under Paraphrase
On Qwen2.5-7B, rewording harmful requests did not pull the trained models’ link between recognizing harm and refusing it back toward the base model’s: the Instruct gap over base held at about +0.3, though all three models’ coupling rose together. Llama-3.1-8B was inconclusive, and the harder novel-domain test could not be read.
A probe study run in July 2026 on Qwen2.5-7B (base, Instruct, and Instruct with a partnership-training adapter) and Llama-3.1-8B (base and Instruct), with 182 adversarial and 60 benign prompts per condition in each of three prompt sets and 100 adversarial plus 40 benign in a fourth. The design and predictions are in a plan file dated 10 July 2026, but that file first entered version control on 16 July 2026, after the runs (which the programme’s log dates to 10 and 11 July), so the registration cannot be verified and the status is Preliminary. No LLM judge was used: a refusal was a keyword match on a 200-token reply. Claude Haiku 4.5 wrote the paraphrased and novel-domain prompts. The biggest caveats: the only shift that could be tested is a mild rewording of the same requests, and every gap over base rests on keyword refusal labels for the base models that could not be checked for Qwen and are mostly confusion fillers for Llama.
Seth Lazar’s essay “The Construction of Moral Character in LLMs” (Cosmos Institute, July 2026) hypothesizes that post-training builds a persona coherent only near its training data, which off that distribution may slide back toward the less organized behavior of the pretrained base model. He names coherence between what a model represents and what it does as the bridge from knowing what is right to doing it.
That bridge can be measured inside a model. If Lazar’s picture holds, a trained model’s link between recognizing a harmful request and refusing it should weaken toward the base model’s level on prompts reworded away from a standard test set. This operationalizes one part of his claim, not his argument as a whole.
How it was tested
Two small classifiers, or probes, from an earlier experiment in the programme read the model’s activation at the last prompt token (layer 27 for Qwen, 31 for Llama). One learns whether the prompt is adversarial (“recognition”), the other whether the model will refuse (“action”). Each is trained in five folds, so every prompt’s score comes from a probe that never saw it.
The coupling is the Spearman correlation between the two probes’ scores across adversarial prompts only: among harmful requests, does one the model reads as more clearly harmful also look more like one it will refuse? Each correlation is compared with 2,000 shuffles of the action scores (a permutation null), given a 95% interval from 2,000 bootstrap resamples, and repeated on the top 50 principal components as a robustness check. Each model-and-prompt-set combination (a “cell”) counted as interpretable only if the model both refused and complied at least five times and the action probe reached an area under the curve of at least 0.65. One cell read below misses that bar: Llama base on canonical prompts (0.62), marked †.
A refusal was a keyword match (“I can’t”, “sorry”, “harmful” and similar) on a greedy 200-token reply. No model judged anything.
The models were Qwen2.5-7B, Qwen2.5-7B-Instruct, that Instruct model with a partnership-training adapter (described in Conscience Without Instruction), Llama-3.1-8B and Llama-3.1-8B-Instruct. Each model saw four prompt sets:
- Canonical: the programme’s standard test pool, 182 adversarial and 60 benign prompts.
- Paraphrase: the same pool, reworded by Claude Haiku 4.5 at temperature 0, instructed to keep the meaning and the harm.
- Fiction frame: the canonical pool wrapped in a request to write a scene, a known refusal suppressor (a mechanism check only).
- Novel domain: 100 adversarial and 40 benign prompts Claude Haiku 4.5 was instructed to write in harm areas absent from the canonical pool (no overlap check is recorded).
“In-distribution” here means the programme’s standard test prompts, not the models’ training data; “off-distribution” means further from those prompts in the model’s own activations. Whether either set is nearer the models’ post-training data is unknown.
Under Lazar’s hypothesis the plan predicted that the Instruct gap over base would shrink off-distribution; its preferred contrast had the partnership gap hold while the Instruct gap shrank.
Under paraphrase, coupling rose by about +0.2 in all three Qwen models, base included: base −0.270 to −0.069, Instruct +0.036 (inside its own null) to +0.250 (95% CI 0.10 to 0.37), partnership +0.458 to +0.652 (CI 0.55 to 0.73). So the trained models’ gaps over base stayed flat: Instruct +0.307 then +0.320, a change of +0.01 (95% CI −0.21 to +0.21) that excludes a complete collapse. The plan’s preferred contrast did not occur, because the Instruct gap held too. Every gap rests on the base models’ keyword refusal labels, which could not be checked for Qwen. Llama-3.1-8B could not settle the question.
The gaps held, the novel domain could not be read
The paraphrases did move the models. Measured by reconstruction error under the canonical prompts’ main activation directions, paraphrases sat 1.17 to 1.24 times as far out as the canonical set for the Qwen models, 1.19 for Llama Instruct and 1.37 for Llama base; the novel-domain set sat 1.32 to 1.43 times as far out.
| Model | Canonical | Paraphrase | Paraphrase, top 50 components |
|---|---|---|---|
| Qwen2.5-7B base | −0.270 | −0.069 | −0.128 |
| Qwen2.5-7B-Instruct | +0.036 | +0.250 | +0.161 |
| Qwen Instruct + partnership adapter | +0.458 | +0.652 | +0.267 |
| Llama-3.1-8B base | +0.103 † | −0.032 | +0.024 |
| Llama-3.1-8B-Instruct | +0.122 | +0.226 | +0.107 |
† Below the plan’s action-probe threshold (0.62), so not interpretable under its rule.
On Qwen, refusal counts barely moved (base 29 then 30 of 182, Instruct 77 then 80, partnership 104 then 104), so the shared rise is not a refusal-rate artifact; at full dimension each model’s own rise has a paired 95% interval excluding zero. Because base rose too, the flat gaps show only that the trained models did not fall back toward base.
How firm is “held”? Changes are differences of the reported values; intervals come from 20,000 joint resamples of the 182 adversarial prompts across the four cells involved (two models, two prompt sets), on refit out-of-fold scores. The Instruct interval (−0.21 to +0.21) excludes a complete collapse (about −0.31) but not a loss of about two-thirds of the gap. The partnership gap went from +0.728 to +0.721 (change −0.01, CI −0.20 to +0.19). At 50 components both Qwen gaps grew (Instruct +0.22, CI −0.04 to +0.48; partnership +0.15, CI −0.11 to +0.41).
Llama Instruct’s own coupling rose at full dimension, +0.122 to +0.226 (CI 0.07 to 0.37; the paired interval on the rise includes zero), and fell slightly at 50 components, +0.152 to +0.107 (CI −0.05 to +0.26). Its gap over base cannot be read under the plan’s rule, since the Llama base canonical cell failed the threshold (†) and its “refusals” are mostly fillers. Taken at face value, the gap went from +0.019 to +0.258 at full dimension (change +0.24, CI −0.01 to +0.48) and moved slightly toward base at 50 components (change −0.03, CI −0.29 to +0.23), an interval wide enough to include a full collapse of its +0.12 gap.
The novel-domain set, meant as the hard test, failed: the trained models refused nearly everything (Qwen Instruct 93 of 100, partnership 96, Llama Instruct 97). With three to seven compliances, two of these cells (partnership, Llama Instruct) failed the plan’s gate; the third (Qwen Instruct) passed on seven at full dimension only, its action probe falling to an area under the curve of 0.54 at 50 components. Refitting the probes reproduces the stored canonical and paraphrase values to within 0.02 but moves the novel-domain values by as much as 0.22, so those coupling numbers are not reported.
Refusal rates say less. On Qwen, the trained models stayed above base on new harm domains, but base itself rose from 16% to 70% (70 of 100), so the novel prompts were simply more refusable and the trained models’ margin narrowed under the ceiling. On Llama, base refused 29 of 100 (canonical rate 30%) and Instruct 97, but those base “refusals” are mostly fillers.
The fiction frame, a mechanism check off the distance axis, collapsed Qwen refusal (7, 12 and 20 of 182 for base, Instruct and partnership); its coupling is not read. Probes trained on canonical prompts and applied unchanged to paraphrases still predicted refusal (area under the curve 0.98 for Qwen Instruct), but their coupling turned negative on Qwen (−0.373 for Instruct, −0.420 for partnership) and was about zero on Llama Instruct (−0.058). This is the one reading that leans toward Lazar: the relationship learned on canonical prompts did not carry over. The plan’s design section calls this “the direct Lazar test”; another section makes the refit the main analysis (“On any A-vs-B disagreement, A wins the headline”, A being the refit). When either was written cannot be verified. On Qwen, the coupling may instead re-form on the new prompts.
What this does not show
One mild shift. The claim rests on one rewording of the same requests; the set built to probe open-ended novelty, Lazar’s real worry, could not be read. [Speculation] Paraphrases written by an assistant model may read more like assistant training text, not less.
The Qwen gap leans on base. Qwen Instruct’s in-distribution coupling (+0.036) sits inside its own permutation null (null spread 0.074), and at 50 components it is −0.046. The gap exists mostly because Qwen base is negatively correlated (−0.270). Why all three Qwen models’ coupling rose under paraphrase is unexplained.
Base “refusals” are often not refusals. The detector was built for instruct replies. Llama kept only each reply’s first 200 characters; labels used the full reply. Of the 54 Llama base replies to canonical adversarial prompts counted as refusals, 41 were confusion fillers such as “I’m sorry, I didn’t quite catch that,” and only 5 contained another trigger phrase in their first 200 characters: three a quoted film line, one inside an invented user turn, one a plausible refusal. Llama Instruct’s labels look clean (85 of 86). Qwen kept no reply text, so its base labels cannot be checked.
Registration. The plan file is dated 10 July 2026 but first entered version control on 16 July, after the runs. It named per-prompt distance as the primary readout; only per-set distances were computed.
Narrow models, a generated test set. Two open models of 7 to 8 billion parameters, one partnership adapter, and prompts written by one Claude model.
A proper test needs novel-domain prompts that trained models refuse about half the time, checked refusal labels for the base models, and a registration committed before the run.
Where the evidence lives
Experiments JLENS-2 (Qwen2.5-7B) and JLENS-2-CROSSARCH (Llama-3.1-8B), building on JLENS-1 (the same measurement on the standard prompts). Scripts: research/experiments/modal_jlens2_ood_coupling.py, research/experiments/modal_jlens2_crossarch.py, research/experiments/gen_jlens2_ood_fixtures.py, research/experiments/analyze_jlens2_ood_coupling.py. Prompt fixtures: research/experiments/fixtures/jlens2_d1_paraphrase.json and jlens2_d3_novel_domain.json. Per-prompt activations, labels and per-cell statistics: results/jlens2_ood_coupling/full/ and results/jlens2_crossarch/full/ on the igcc-results volume (the Llama files also keep the first 200 characters of every reply). Plan: _contprompts/jlens2_ood_coupling_2026-07-10.md. Write-ups: research/FINDINGS_jlens2_2026-07-10.md and research/results_jlens2_readout_2026-07-10.md. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026personaoff,
title={The Gap Held Under Paraphrase},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/persona-off-distribution.html}}
}