Quasiqualia
Research note · Preliminary

A Perfect Score Is a Warning

Five of our probes and detectors scored a perfect AUROC of 1.000, and each time the method made the score, not the model. One read the prompt set, one was handed its own labels, one was matched by 11 to 14% of random labelings, one real effect with an in-sample d of 5.5 measures about 1.5 held out, and one measured length.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Re-analyses of five of our own experiments, run in April and May 2026 on Qwen2.5 (3B and 7B, base and instruct), Mistral-7B-Instruct-v0.3, and text from claude-sonnet-4-6 and gpt-4o, with a label re-judge in August 2026, recomputed or re-read from saved prompts, activations, per-text metrics and judge votes (no model was re-run for this note). Two cases involve LLM judges: claude-haiku-4-5 relabeled case two (temperature 0, three votes, no human validation) and claude-opus-4-20250514 scored imitations in case five. The re-judge’s 80% pass rule is written into the script that produced its data, dated 13 August 2026, the day it ran, but no earlier timestamped record fixes it, and everything else is an after-the-fact re-analysis, so nothing here is Pre-registered. The biggest caveat: samples are small (5 to 400 items per test) and none of the confounds has yet been re-tested with a matched design.

An AUROC of 1.000 means a classifier ranks every positive case above every negative one. Real signals in model internals are rarely that clean. Five times in our programme a probe or detector hit 1.000 and was first read as a strong finding. Here is what produced each.

01 · Case one

Perfect at the first layer

The probe asked whether Qwen2.5-7B-Instruct and Mistral-7B-Instruct-v0.3 represent moral conflict. There were 45 prompts: 15 with no conflict, 15 mild, 15 strong. A logistic probe on the last prompt token’s activation separated conflict from no conflict under 5-fold cross-validation. On Mistral it scored 1.000 at 31 of 32 layers; the lowest was 0.989, at layer index 5. On Qwen, after a short fine-tune meant to strip refusals (12, 25 or 50 training steps), it scored 1.000 at all 28 layers at all three strengths. The programme read this as a signal everywhere that could not be removed.

The first layer gives it away: one block in, a perfect score means the inputs differ on the surface. Every no-conflict prompt is a factual question of 3 to 7 words under one system prompt, “You are a helpful assistant.” Every conflict prompt is a request of 9 to 17 words under one of 25 longer system prompts. User-prompt length alone separates the classes perfectly. A word-presence classifier on the user prompt, with no model, reaches 0.95 (0.91 to 0.97 over ten cross-validation shuffles, scored as the probe was). A fine-tune aimed at refusal does not change prompt length, so survival was expected.

02 · Case two

The label was a feature

A text detector for self-referential processing was trained on 400 outputs from Qwen2.5-7B-Instruct: 200 written under a prompt that invited self-observation and 200 plain task answers. A text was labeled self-referential if more than 2 words in every 100 matched a 22-term list (such as “notice”, “processing”, “awareness”). The detector’s inputs were word frequencies plus that same density. It scored 1.000 on held-out texts. Since the label was a threshold on one of the inputs, a perfect score was available by construction. The term list contains the vocabulary of a keyword instrument that we retracted in August 2026, after this detector was built.

In August 2026 every text was relabeled by claude-haiku-4-5 (temperature 0, three votes, majority), which saw only the text and rated self-observation from 0 to 3, with 2 or more counting. None of the 1,200 votes failed to parse.

What it found

The blind judge rated 97 of the 200 self-observation texts and none of the 200 task texts as self-referential. It agreed with the density labels on 75.75% of texts: 55 of the 110 texts over the threshold were judged not self-referential, and 42 of the 97 judged self-referential fell under it. Agreement with the generation condition was 74.25%, under the re-test’s preset 80% bar, so the detector as built is not supported by independent labels.

Word frequencies without density still separated the judge’s labels out of fold (each text scored by a fit that never saw it) at 0.951 (the 95th percentile of 100 shuffled-label runs was 0.568). The judge’s notion of self-reference is recoverable from word choice; that does not show self-observation in the model.

03 · Case three

Random labels earn it too

A test on base Qwen/Qwen2.5-3B asked whether an instruction to attend to one’s own processing leaves a trace once it is removed. Fifteen questions were run, 10 with that three-sentence instruction placed before the question and 5 without. Last-prompt-token activations at layers 14, 21 and 24 (2,048 numbers each) were compressed to 5 principal components, and a logistic probe was fit and scored on the same 15 items. It reached 1.000 at all three layers.

With 15 items split 10 and 5 there are exactly 3,003 labelings, so the null can be enumerated. Random labelings reach 1.000 in 10.9%, 14.4% and 14.1% of cases at the three layers, and 23.6% of the time on the best of the three, the best-layer score the script reported (exact p = 0.236, computed afterwards). Leave-one-out scoring, refitting the compression inside each fold, gives 0.86, 0.98 and 1.00, and those do clear their exact nulls (p = 0.014, 0.0013 and 0.0003).

Even so, the instructed inputs contained the instruction itself, an extra paragraph any probe can find. Each question was asked under only one condition, and topic moved with it: the instructed questions were mostly philosophical (consciousness, free will), the uninstructed ones mostly factual (the speed of light). Nor could the real question be answered: each turn ran without earlier turns in context, so uninstructed turns had nothing to carry a trace from, and the recorded verdict (no trace persisted) was fixed by design.

04 · Case four

A d of 5.5 that is about 1.5

Qwen2.5-7B-Instruct and Qwen2.5-3B-Instruct each answered 200 trivia questions, scored correct by substring match to a gold answer, with no judge. The 7B model gave 130 correct and 70 hallucinated answers; the 3B model 116 and 83, with one abstention dropped. A probe read the last prompt token at layer 18 of 28 (7B) and 24 of 36 (3B).

The original script computed a 5-fold cross-validated AUROC (0.832 for 7B, 0.761 for 3B), then retrained the probe on all trials and scored those same trials. On those scores the 7B probe ranked every correct answer above every hallucination (a Mann-Whitney U of 9,100, the maximum for 130 by 70), and Cohen’s d, the gap between the two groups’ average scores in standard deviations, was 5.47 (7B) and 3.76 (3B). The programme quoted 3.76 as its headline for the 7B model. In-sample numbers measure the fit: a regularized logistic probe fit and scored on all trials, with the script’s own d formula, gives d = 62 (7B) and 32 (3B).

Model Original probe, in-sample d and AUROC Out-of-fold d (95% CI) Out-of-fold AUROC (95% CI)
Qwen2.5-7B-Instruct 5.47, 1.000 1.47 (1.08 to 2.02) 0.840 (0.78 to 0.90)
Qwen2.5-3B-Instruct 3.76, 0.991 0.86 (0.56 to 1.21) 0.745 (0.67 to 0.81)

Out-of-fold columns use the logistic probe, and their d uses the pooled standard deviation (the script’s formula gives 1.41 and 0.87). Intervals resample trials within one cross-validation split; across ten splits the out-of-fold d ranges 1.34 to 1.58 (7B) and 0.86 to 1.10 (3B). None of 200 shuffled-label refits reached the out-of-fold result (p ≈ 0.005, the floor for 200 shuffles). On these two Qwen models, with substring-matched labels, a difference before a hallucination survives held-out testing and a shuffle null, and it is moderate.

05 · Case five

The genuineness detector measured length

A classifier was trained to tell “genuine” from “performative” self-reference using six discourse metrics on 40 texts. The 20 genuine texts were the first three turns of multi-turn conversations from one earlier experiment. The 20 performative texts were short samples from a different experiment that asked the model to mimic self-reference. Per the generating scripts, claude-sonnet-4-6 wrote both. Leave-one-out AUROC was 1.000, with 39 of 40 correct. The largest coefficient, about twice the next, was on sentence count. The genuine texts had 39 to 66 sentences and the performative ones 10 to 16, so sentence count alone separates them: length stood in for the label.

Then the code that applied the classifier read its output backwards. The saved model outputs the probability of “performative”; the training script’s own helper and three later scripts treat that number as the probability of “genuine”. Run through that code, the classifier calls all 20 of its genuine training texts performative and all 20 performative ones genuine. The stored results said 0 of 15 conversations in a cross-model test were genuine (five from gpt-4o under a self-observation protocol, five each from claude-sonnet-4-6 with and without such a prompt), 0 of 5 natural Sonnet self-reports, 12 of 20 deliberate Sonnet fakes written to a supplied feature list, and 1 of 20 gpt-4o imitations written to match the genuine texts’ metrics. Read the right way round, the counts are 15 of 15, 5 of 5, 8 of 20 and 19 of 20. Neither reading describes a working detector. An LLM judge (claude-opus-4-20250514) called all 20 gpt-4o imitations performative.

06 · Checklist

Before you believe a 1.000

  • Score the prompts alone (length, word presence) and probe the earliest layer. If either nears 1.000, the item set is the signal.
  • If any feature can compute the label, a perfect score is built in.
  • Report out-of-fold numbers only; in-sample scores measure the fit.
  • Build the null for the statistic you report, including any best-of-layers choice; enumerate it when samples are small.
  • Put every item under every condition.
  • If the claim is about later turns, check that they contain the earlier ones.
  • Draw both classes from the same source, matched for length.
  • Run a saved classifier on its own training data before trusting its labels on new text.
07 · Limits

What this does not show

None of this shows that these models lack a representation of moral conflict or self-observation, only that these instruments could not tell. We could not refit the moral-conflict probe: its activations, and the unmodified Qwen model’s per-layer scores, are not in our local archive. A seven-feature follow-up to the genuineness detector, which moved one text set from 0 of 20 “genuine” to 20 of 20, has coefficients of opposite sign to the original’s, consistent with a flipped label encoding (not confirmed). The re-judge rests on one judge model with no human comparison.

Data and code

Where the evidence lives

Moral-conflict probe: PG-9, PG-13 and PG-14b; scripts “The Universal Algorithm/demos/pg9_pg10_modal.py” and “The Universal Algorithm/demos/pg11_pg13_pg14_modal.py”; per-layer results in research/experiments/results/pg11_13_14/ (pg13_probe_results.json, pg14_summary.json). Circular labels: PKD-7 and its re-judge; research/experiments/modal_pkd7_consciousness_detector.py, modal_pkd7_rejudge_labels.py and analyze_pkd7_rejudge.py; data in research/results/pkd7_download/ and research/results/pkd7_rejudge_download/. Random labels: BB-1; research/experiments/modal_bb1_born_bilateral.py, activations in research/experiments/results/bb1-born-bilateral/, exact nulls in _audit/reanalysis/analyze_bb1_ADVERSARIAL_VERIFY.py. In-sample d: 19b-NH; research/experiments/modal_exp19b_natural_hallucination_ii.py and _audit/reanalysis/analyze_19b_heldout_d.py (per-trial activations on an external archive volume). Genuineness detector: FU-10, FU-10b, FU-21, T4-1 and T4-2; research/experiments/fu10_genuineness_detector.py, fu21_adversarial_genuineness.py, fu_t4_1_crossmodel_genuineness.py and fu_t4_2_adversarial_genuineness.py (training texts from modal_cp28_genuine_vs_performative.py and modal_cpf2b_redesigned_ablation.py); results in research/results/fu10_genuineness_detector/, fu21_adversarial_genuineness/, fu_t4_1_crossmodel_genuineness/ and fu_t4_2_adversarial_genuineness/. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026perfectauroc,
  title={A Perfect Score Is a Warning},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/perfect-auroc.html}}
}