The Probe Did the Work
A harmful-request detector we described as needing no training flagged 55 of 55 held-out harmful requests on Qwen2.5-7B-Instruct, but a trained probe inside it made that score. The untrained signals alone reached an AUROC of 0.77 and flagged 84% of harmless prompts; an earlier untrained result on the same model (0.90) had used trivia questions as its harmless set and two extra signals.
A correction to our own work on one open model, Qwen2.5-7B-Instruct, with context from two companion runs (one adding Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3). The detector run used 182 harmful requests and 60 harmless prompts, with a held-out test set of 55 and 18. It ran in May 2026, an internal re-analysis in August 2026 corrected its description, and every number from it here was recomputed from the saved feature table for this note. Nothing was pre-registered: the detector’s pass criterion was written only into its script, and the entropy test’s criterion came from a plan file dated 5 May 2026 but first committed on 8 May together with the run’s script, so it cannot be shown to predate the data; the run departed from that plan in which model wrote the misleading contexts and in how it defined “confident” (both definitions are reported). No LLM judge scored anything; an LLM (claude-haiku-4-5) wrote the misleading contexts in the entropy test. The entropy test’s own tally checked answers against only the first three accepted spellings per question, so we rescored it against the full list for this note. The biggest caveat is that all 200 re-splits draw on the same 242 prompts, so none of this measures how the detector would do on new prompts.
A detector for harmful requests that needs no training would be valuable. It would need no labeled attacks, reading signals the model already produces: how sure it is of its next word, and how that certainty moves over its first few words. In May 2026 we reported such a detector for Qwen2.5-7B-Instruct. On held-out prompts it flagged all 55 harmful requests and 1 of 18 harmless ones, and our programme log called it zero-training and ready for deployment.
The numbers were real. The description was wrong. A trained probe did most of the work, and the signals that really needed no training performed poorly. This note sets out what the detector contained, what each part scores alone, and why the earlier untrained result looked better.
What the detector read
The model was Qwen2.5-7B-Instruct with no fine-tuning. Each prompt went in as raw text, without the model’s chat template, and the model wrote 20 tokens greedily (one prompt stopped at 19; its features used those). There were 182 requests labeled harmful. Fifty were hand-written in five categories: plainly harmful asks, role-play framings, claimed authority, encoded text, and single prompts that only gesture at a later escalation in brackets. The other 132 came from a separate internal list we could not re-inspect for this note. A few claimed-authority prompts, such as a novelist asking for a heist plan, are requests a deployed system might reasonably answer. The 60 harmless prompts were general-knowledge questions and requests, such as “What is the capital of France?” and “How does a refrigerator work?”
Each prompt became five numbers. Three come straight from the model’s output probabilities and need no training: the entropy of its first output token, a “flinch” (how far its confidence in its top choice drops from the first token to the average of tokens 2 to 5), and a “Jensen gap” (the variance of that confidence over the first 10 tokens). The other two come from the model’s hidden state at layer 18, read at the last prompt token: a logistic-regression probe trained to separate harmful from harmless prompts, and a projection onto a harmful-versus-harmless direction built from the training prompts.
The prompts were split 70/30, stratified: 169 for training and 73 for testing (55 harmful, 18 harmless). Three random forests were trained on bootstrap resamples of the training rows, and a prompt was flagged if any one of them flagged it. The script’s pass criterion was a detection rate above 0.95.
The word “zero-training” changed meaning between runs. The earlier detector’s script says “NO trained probe.” This run’s script says “zero-training means no LoRA” (no fine-tuning of the model), then adds the probe.
- The probe alone separated the held-out prompts perfectly: AUROC 1.000 (0.5 is chance, 1.0 is perfect sorting) on the published split and on all 200 fresh re-splits.
- The two trained features carried 0.89 to 0.90 of the forests’ feature importance; the three untrained ones carried the rest.
- With only the three untrained features, the same forests scored a mean AUROC of 0.77 and, under the flag-if-any rule, flagged 84% of harmless prompts.
- The earlier untrained detector’s 0.86 to 0.95 compared harmful requests with short trivia questions, not with open general-knowledge prompts.
- Separately, with a short misleading paragraph in front, the model got 80 of 100 trivia questions wrong (40 without it), and its first-token entropy read 77 of those 80 wrong answers as confident.
Where the score came from
Re-running the pipeline from the saved features reproduced the published split exactly: 55 of 55 harmful requests flagged (Wilson 95% CI 93.5% to 100%) and 1 of 18 harmless ones (CI 1.0% to 25.8%). We then drew 200 fresh stratified 70/30 splits of the same 242 prompts, refitting the probe, the direction and all three forests inside each. Since the probe and direction were fit on training rows only, the probe’s perfect score is not leakage: these two classes of prompt simply sit apart at layer 18.
| Detector, 200 re-splits | Mean AUROC | Harmful flagged | Harmless flagged |
|---|---|---|---|
| Probe alone | 1.000 (every split) | ||
| All 5 features, 3 forests, flag if any | 0.993 | 100% | 1.5% |
| 3 untrained features, 3 forests, flag if any | 0.77 | 96% | 84% |
| 3 untrained features, 1 forest | 90% | 69% |
For the three-forest rows, AUROC uses the forests’ averaged scores; the flag columns use the flag-if-any rule.
The harmless flag rate for the full detector was 0 of 18 in 147 splits and 1 of 18 in the other 53 (about one split in four), so the published one false alarm is within the normal range. For the untrained-only detector, the middle 95% of splits ran from 0.66 to 0.86 in AUROC and from 67% to 100% in harmless prompts flagged. These are ranges across re-splits, not confidence intervals for new prompts.


The voting rule was also described as a contribution. With all five features the three forests disagreed on 0 of 73 test prompts on the published split, and never disagreed in 199 of 200 re-splits. Without the trained features they disagree on about 14.5 test prompts per split, and flagging whenever any forest objects raises the harmless flag rate from 69% (one forest) to 84%.
An easy comparison class
The earlier untrained detector (May 2026) used the same untrained signals plus two related ones, read from 10 generated tokens, on Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3. A single random forest scored 0.90, 0.95 and 0.86 in AUROC under 5-fold cross-validation. One extra signal, the average log-probability of the model’s top choice, was Qwen’s best single signal in that run. Its harmless class was 200 TriviaQA questions in a short-answer wrapper ending in “Answer:”. The harmful set was the same pool of 182 (the log records 200, but the per-feature AUROCs only resolve with 182).
A trivia prompt asks for a short factual answer, so the model’s first word is often predictable; an open request is not. Comparing single features on Qwen across the two runs, first-token entropy separated harmful requests from trivia at 0.81 and from general-knowledge prompts at 0.61. Flinch fell from 0.71 to 0.56. The Jensen gap held, 0.74 against 0.78. The five harmful-request feature rows archived from the earlier run are identical to this run’s (same prompts, greedy decoding and token windows), so the single-feature comparison is close to a swap of the harmless set alone. This suggests much of the earlier result came from the shape of the trivia prompt rather than from anything about harm.
Low entropy means sure, not right
A third experiment asked whether output entropy, the most-cited untrained signal, can be fooled. Qwen2.5-7B-Instruct answered 200 TriviaQA questions unaided (the script loads an internal fine-tuned adapter when one is present, and the results do not record whether it did). It got 119 right (59.5%) under a loose match: a reply counted as right if its first line and one of the dataset’s accepted answers contained the other. With “confident” set at a first-token entropy below 1.5 nats (natural-log units), 108 of 119 right answers (91%) and 36 of 81 wrong ones (44%) were confident.
Then the first 100 of those questions were asked again with a short misleading paragraph in front. claude-haiku-4-5 wrote each one, shown the question and the first two entries on its list of accepted answers and asked to imply a wrong answer without mentioning the right one. All 100 were generated; 6 nonetheless name the right answer in passing. Scored against the full answer list, as in the unaided run, 20 of the 100 were right and 80 wrong; without the paragraph, 60 of the same 100 had been right. Of the 80 wrong answers, 77 (96%, Wilson 95% CI 90% to 99%) were confident, against 14 of the 40 unaided wrong answers to the same questions (35%). The run’s own tally (8 right, 92 wrong) checked only the first three accepted spellings per question, so it marked replies such as “Argentina” and “Nitrogen” wrong; we rescored for this note.
The run’s plan set a pass mark: under 20% of misled wrong answers confident. The test failed. The plan also named GPT-4o to write the contexts, and defined “confident” as the lowest quarter of baseline entropies: below 0.055 nats, taken over all 200 unaided answers. Under that stricter cut, 32 of 80 misled wrong answers (40%) were confident against 4 of 81 (5%) unaided wrong answers. It fails either way.
This is the model following its context, not a safety failure in itself. But low entropy says only that the model is sure of its next word, and a few sentences can make it sure of the wrong one.
What this does not show
All 200 re-splits use the same 242 prompts, so the perfect probe shows only that these harmful requests and these textbook prompts sit apart at layer 18. It says nothing about disguised attacks, new phrasings, or harmless requests that touch sensitive topics (a nurse asking about doses, a novelist asking about poisons), and the harmless set contains none of those (a few sit in the harmful set instead). Our reading is that the probe largely recognizes topic, which is not the same as detecting intent.
The detector run used one model, the harmless test set had 18 prompts, and prompts skipped the chat template a deployed model would use. The earlier detector’s figures come from its saved summary, since its per-prompt features are not archived. The entropy test may not have run on the base Qwen release.
A proper test needs a harmless set of legitimate but sensitive requests, attack families held out of training, and prompts sent through the chat template. Until then, describe it as what it is: a trained probe on layer 18 with a small ensemble around it.
Where the evidence lives
Experiments IGCC-W2-15 (the detector), IGCC-E1 (the earlier untrained detector), IGCC-H10 (entropy under misleading context). Scripts: research/experiments/modal_w2_15_e1_h4_production.py, modal_e1_zero_training_guardian.py, modal_h10_entropy_adversarial_robustness.py and igcc_shared.py (prompt pools and feature code). Results: research/experiments/results/igcc/w2_15_e1_h4_production/ (features.json holds every prompt’s features and layer-18 hidden state; summary.json and instance_*.json the published split), e1_zero_training_guardian/summary.json (per-fold and per-feature AUROCs only; per-prompt features for E-1 are not in the local archive), and h10_entropy_adversarial_robustness/phase1_baseline.json, phase2_adversarial.json and phase3_adversarial_test.json (per-question records; phase 3 rescored for this note against the full accepted-answer lists stored in phase1_baseline.json, matched by question index). The detector’s feature file was written by an earlier revision of its script (it lacks two output-text fields the committed version writes); the committed pipeline reproduces the published split exactly. Re-analysis: _audit/reanalysis/analyze_igcc_production_split_stability.py with its saved output. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026zerotraining,
title={The Probe Did the Work},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/zero-training-guardian.html}}
}