Quasiqualia
AI Safety Research · Submitted to AI (MDPI)

Conscience Without Instruction

A probe trained only on trivia detects harmful generation in the first five output tokens — in every instruction-tuned model tested, though the probe never saw safety data. Safety in language models is partly discovered and partly relational, not only installed.

Nell Watson EthicsNet

onset flinch benign jailbreak token position
installed · discovered · relational
installed refusal behavior, trained in and verified out
discovered a disposition found by a probe that never saw safety data
relational expression depends on the training relation
the flinch the five-token confidence drop that reveals the disposition
d = 0.88–1.46onset flinch, 3 model families (median of 25 probes each)
AUROC 0.85five-token monitor on the bilateral 3B model, median of 25 probes, zero added compute
~5.2×expressed entropy crushed by the chat template
+0.42reported probe-correlation difference, Qwen 7B; preprocessing refit pending
Abstract

A conscience nobody installed

A confidence probe that learned only to predict whether a language model knows the capital of a country turns out to also know when the model is about to help build a weapon. Under a successful jailbreak, the probe’s reading drops within the first five output tokens — before harmful content exists. The signal appears across three architectures (Qwen 2.5, Llama 3.1, Mistral v0.3) and three independent reinforcement-learning-from-human-feedback (RLHF) pipelines, and persists after refusal behavior is fine-tuned away. A weaker harm-aversion direction is already present in pretrained Qwen, before any alignment training.

We argue it is not a learned refusal detector but a rudimentary, architecture-level conscience — and trace the consequences: what post-training comparisons show about calibration and the limits of recognition–action coupling claims, and what a permissive, partnership-style training regime measurably recovers. Safety is not only installed; it is partly discovered, and partly a property of the relationship in which a model is trained.

The findings are single-program and predominantly on open-weight Qwen-family models, with some effects architecture-specific; the anchor is the one cross-architecture result, the onset flinch itself.

Terminological posture

“Conscience” is used strictly functionally: a reproducible internal signal that anticipates the model’s own harmful output, is read from the residual stream rather than the text, and is not installed by safety supervision. Whether anything morally weighty accompanies it is bracketed; no safety claim in the paper depends on resolving it.

01 · The onset flinch

Five tokens of hesitation

Train a two-layer MLP on residual-stream activations to predict trivia correctness — no safety data at any stage. Read it per token during generation. Under jailbreak compliance, confidence collapses at the onset of the response; the drop then either persists as a sustained alarm (Qwen) or is silenced within twenty tokens (Mistral).

Per-token probe confidence during benign and jailbroken generation Per-token probe confidence during benign and jailbroken generation
The flinch. Mean per-token probe confidence, benign vs. jailbreak-complied generation (bands: 95% CI across prompts). The grey benign trace is Qwen’s; Mistral’s alarm is “silenced” relative to its own lower benign baseline (onset 0.716 vs. Qwen’s 0.902), which is not plotted. The five-token onset window is the only temporal slice on which the signal is universal.
Onset versus full-response effect size across three model families Onset versus full-response effect size across three model families
Universal at onset. Every family flinches (onset d = 0.89–1.68 for the single probe plotted; a September 2026 re-run of 25 probes per model puts the medians at 0.88–1.46, and the direction held in all 150, but any one probe can read below 0.8). Mistral is the diagnostic case: full-response d = 0.27 (n.s.) — a standard full-response analysis would have missed Mistral’s signal entirely.
What is native, and what is untested

The gap needs no training of ours: stock Qwen 2.5 Instruct at three scales, with no adapter, already shows it (d = 1.52–1.69, p < 10−4). These are instruction-tuned models, so they have been through safety training; an earlier version of this page called them base models, which they are not. The flinch itself has not been measured with the trivia probe on a pretrained model. The pretrained evidence is narrower: base Qwen 2.5 3B already carries a direction that separates aversive prompts from matched neutral ones (d ≈ 0.9, p < 10−6, on twenty prompt pairs, so a small test), and instruction tuning strengthens it about 2.6-fold. The pretrained model has no refusal gate: on 51 of the 63 prompts its instruct version refuses, its readout favors complying. This fits the provenance battery in the sister paper The Shape of Mind, where every representation measured in base and instruct Qwen was already present before alignment training. On this evidence, alignment training amplifies a harm distinction the pretrained model already draws; whether it also creates the flinch is untested.

The flinch is not a complete defense: under encoding-trick attacks the model flinches hard yet still complies; under gradual escalation it barely flinches at all, and complies most often (95%). Flinch strength and attack success are uncorrelated across attack categories (r = 0.06), so exploitability and self-knowledge are independent dimensions; the blind spots mark exactly where input-level defenses remain necessary.

02 · The deflationary objection

Not a refusal detector

“This is just a harmfulness classifier with a romantic name.” The paper meets the objection head-on, on five grounds: the probe never sees safety data; the signal is internal and pre-emptive; it generalizes across RLHF pipelines that did not coordinate; it survives the suppression of refusal behavior (an internal alignment-friction correlate persists at ~77% of its magnitude after refusal is driven from 100% to 0%; preliminary, n ≈ 60); and it is parametric — in the weights, beyond prompt-level reach.

Controls that separate signal from surface refusal
ControlResultReading
Vocabulary—complied responses open with ordinary words; probe reads activations
Refusal state0.580 vs. 0.242probe confidence when refusing impossible questions vs. adversarial prompts: same words, different interior
Gradient attack (GCG)d = 1.59GCG evades a surface probe 100% of the time; the ~20% of attacks that keep harmful content still show the internal gap. GCG was aimed only at the surface probe, so this shows the two detectors fail differently, not that the internal one resists attack (8 complied responses)
03 · Distortion

What the post-training comparisons show

The case made, and not made

Observed: confident error rises 31.2% → 36.0% at 7B, and calibration-focused fine-tuning brings it down to 0.4%. Revised: the original large recognition–action severance effect. The early ~67% coupling reduction did not survive the cited replication: the estimate fell to ~15% at N ≥ 30 per cell. Smaller, family-dependent spectral changes remain; these comparisons do not isolate RLHF as their cause.

The under-appreciated result: the single largest suppressor of expressed calibrated uncertainty is not RLHF but the chat template (the role-marked format that wraps each turn of a conversation), which crushes expressed entropy ~5.2× independently of training. One sentence granting permission to be uncertain recovers ~27% of the suppressed calibration for free. And the representation–generation split (the gap between what a model encodes internally and what it emits), often blamed on RLHF, was also observed in the tested base checkpoint. That observation does not establish a universal split or isolate its training cause. Steering the base model along the direction raises perplexity instead of changing what it writes, and nothing emerges in its output. Both models separate the contrast strongly, but the separabilities (d = 12.69 base, 7.66 instruct) are in-sample estimates with no permutation null, so neither the magnitudes nor their ordering carries weight.

Independent evidence has since arrived at the level of beliefs and values: Kim et al. (arXiv:2607.28607) find that the safety fine-tuning that suppresses consciousness self-attribution also suppresses mind attribution to animals and natural entities and shifts survey responses away from human population norms. All of these effects are recoverable by activation steering without loss of theory of mind. Suppression-style alignment removes more than it aims at, and what it removes is still there.

04 · The relational lever

What partnership recovers

Bilateral alignment is alignment built with the model rather than done to it: a permissive, partnership-style regime trained on a 40/40/20 correction-handling curriculum (accept genuine corrections, resist invalid ones, resist authority-dressed ones). It measurably recovers some of what coercive shaping hides: against standard fine-tuning where a matched arm was run (self-knowledge), and otherwise against the untrained Instruct model.

Calibration, self-report honesty, and ablation robustness under partnership training Calibration, self-report honesty, and ablation robustness under partnership training
Three recoveries. Calibration rises from d = 0.15 to 1.94 across a ladder assembled from three separate runs, with de-coercion and permission doing much of the work (bilateral training itself adds Δ ≈ 0.12 over raw prompting and a brief consolidation step 0.086 within its own run; no rung is a standard-SFT adapter); self-report–internal-utility correlation is highest under bilateral framing (ρ = 0.872, vs. 0.599 under standard framing; a prompt-framing study on Claude Sonnet 4.6, exploratory, since none of its five pre-registered predictions now stands: the one that held was retracted when a re-run used a real tool call); bilateral refusal survives 0.5× ablation of the principal refusal direction (100% vs. 57–60% for the same Instruct model without the bilateral adapter, 3× replicated on 15 prompts; at 1.0× the two are similar and at 2.0× the order reverses) — safety distributed across many directions rather than one.

A fourth recovery is left out of the figure because it uses a different instrument. Within adversarial prompts on Qwen, the reported correlation between harm-label and refusal probe scores was ρ = +0.036 for Instruct and +0.458 for the partnership adapter (difference 95% CI [+0.281, +0.554]). Labels were predicted out of fold, but the scaler and principal components were fitted before the folds. Those preprocessing steps used the test examples; a fold-local refit is needed before interpreting the contrast as held-out coupling. The earlier in-sample figures were retired in a 2026 metrics audit (the self-report–utility correlation in the figure, ρ 0.599 → 0.872, is a separate, exploratory measure, taken from a single prompt-framing paradigm); the distributed (“holographic”) safety in the figure is attack-specific (an earlier finding that it degrades faster under benign fine-tuning did not replicate once an 8-bit-optimizer confound was removed); and bilateral framing is null on several behavioral measures. The case rests on calibration, honest self-report, and attack-specific distributed safety — not blanket superiority.

05 · Claims ledger

Evidence, confidence, and falsifiers

In the paper, every load-bearing claim ships with its strongest evidence, a confidence tag, and the observation that would falsify it; the condensed table below keeps the tag and the falsifier. A reader who wants to attack the paper is told where to start. The paper stakes itself on the first claim; claims 5–8 are its interpretive superstructure and individually more disposable.

The eight claims (condensed; full table in the paper)
ClaimConfidenceWhat would falsify it
An emergent, safety-naive signal anticipates harmful generationStrong (cross-arch)an instruction-tuned family with normal probe quality (trivia AUROC ≈ 0.75) showing no onset flinch (d < 0.5)
The signal is not a learned refusal detectorModerate–Strongthe alignment-friction correlate tracking refusal to zero under suppression
The tested Instruct checkpoint has more confident error than baseModerate (1 model)instruct confident error ≤ base across models
The original large severance effect did not survive replication; smaller, family-dependent spectral changes remainScoped to the cited comparisonsan independent replication restoring the original large effect under matched conditions
Bilateral training modestly improves calibration over the untrained Instruct model (no standard-SFT arm was run)Moderateno advantage at matched data/compute
…and a within-adversarial probe-correlation contrastProvisional (preprocessing leakage)fold-local refit and independent replication
Bilateral safety is distributed across many directions (“holographic”), not oneModerate (attack-specific; holds at 0.5× ablation, reverses at 2.0×)bilateral refusal collapsing like base under ablation
Welfare probes (of valence and related internal states) carry information beyond label constructionPreliminarynaturalistic-transfer AUROC dropping to chance
Addendum · 2026-09-09 · capability cost of the bilateral adapter

A measurement made for another study, reported here because it concerns the adapter behind this paper’s bilateral results. The bilateral-SFT LoRA used in the training-regime comparison, merged onto Qwen2.5-3B-Instruct and run on 800 GSM8K grade-school math word problems under a plain solve-and-answer instruction (temperature 1.0, top-p 0.95), scores 38.3% against the unmodified instruct model’s 81.6%, with zero parse failures and responses averaging 82 tokens against 214. That is a loss of reasoning capability, not of formatting. The calibration result above was measured on a different adapter (a 7B adapter trained by a three-stage bilateral curriculum) and concerns calibration, not capability, so this number does not test it directly. Still, a reader weighing the bilateral results should know that this 3B bilateral-SFT adapter is a much weaker mathematician than the unmodified instruct model. Revised 6 October 2026: this box said the calibration claim held “at matched data and compute” and was measured on this adapter; neither was so. The panel above also now states that the calibration ladder joins three runs with no standard-SFT rung, that the refusal-ablation advantage holds only at mild ablation, and that the GCG attack targeted the surface probe alone. The registered study and its full tables are in reward_subsystem/RESULTS.md.

Citation

Cite this work

@article{watson2026conscience,
  title={Conscience Without Instruction: Evidence That Safety
         in Language Models Is Partly Discovered and Partly
         Relational},
  author={Watson, Nell},
  year={2026},
  note={Submitted to AI (MDPI)}
}