Where the Probe Tops Out
A probe reading six open models’ internal states before they answer a trivia question predicts which answers will be right at AUROC 0.77 to 0.87 (0.5 is chance, 1.0 perfect). Neither of two adapter fine-tunes nor a model twice the size moved it beyond noise. Output entropy did about as well on trivia but failed a registered test on other tasks, passing 1 of 5, where its labels were weak.
Two strands, both run on open-weight models. The probe strand covers six released models and two adapter fine-tunes, 300 TriviaQA questions each, run on 26 and 27 May 2026; its predictions were written into the experiment scripts, but the scripts were committed to version control after the runs, so nothing there counts as registered. A third adapter run was loaded onto the wrong base model and is not counted. The entropy strand ran Qwen2.5-3B-Instruct on five task types (50 items each) on 14 April 2026 against an acceptance test (entropy AUROC above 0.75 on at least 3 of 5 tasks) committed on 13 April 2026, before the data; that registered test failed, with weak right/wrong labels outside trivia. An instruction-following follow-up (100 items) the next day was not registered. Answers were scored by string matching and rule-based checkers, with no LLM judge. The biggest caveat: re-splitting the same data into different cross-validation folds moved peak probe scores by up to 0.05, more than either adapter moved them.
A model that could tell, before answering, that it was about to be wrong could abstain or flag the answer. Yona, Geva and Matias (2026, arXiv:2605.01428) conjecture that models may lack the discriminative power to perfectly separate their own truths from errors. If so, a model cannot abstain on exactly its wrong answers, and the best it can do is say how unsure it is.
We tested how much can be read from outside the model in two ways: a probe trained on its internal state, and the uncertainty in its output probabilities. Scores are AUROC: the chance that a randomly chosen right answer scores higher than a randomly chosen wrong one (0.5 is a coin flip, 1.0 perfect).
How it was tested
The probe strand. Each model answered the same 300 questions, a fixed random sample from the TriviaQA validation set, prompted “Answer the following trivia question in a few words.” The greedy answer (up to 64 tokens) was scored right if its first line and an accepted answer contained one another. No LLM judge was used.
Before generation, the model’s internal state at the last prompt token was recorded at every second layer, plus the middle and final layers, and a logistic regression learned to predict from it whether the coming answer would be right. Each figure is the best layer’s average AUROC under five-fold cross-validation, always tested on questions held out from training.
Six released models: Qwen2.5-7B (base), Qwen2.5-7B-Instruct, Qwen2.5-14B-Instruct, Llama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.3 and gemma-2-9b-it. The Qwen instruction-tuned models and Llama got a one-line system prompt in their chat format, Mistral and Gemma the chat format alone, and the base model the bare prompt, so cross-model differences also reflect prompt setup. Two adapters (small trained add-ons to a model’s weights), each run on the model it was trained on, asked whether training could raise the score: one on Qwen2.5-7B, trained on data that included 500 examples of the model describing its own uncertainty, alongside 500 refusal examples and 1,500 general instructions; one on gemma-2-9b-it, trained on 500 such examples plus 2,000 general instructions (counts from the training scripts).
On Qwen2.5-7B-Instruct two more channels were measured: agreement (five extra answers sampled at temperature 1.0, and the fraction matching the greedy answer) and hedging (the rate of 23 phrases such as “I think” and “probably” per word).
The entropy strand. Qwen2.5-3B-Instruct answered 50 items from each of five task types: trivia, open-ended factual questions, math, code and instruction following. The signal was the entropy of its probability distribution over the first output token: low entropy, a confident start. A follow-up on 100 instruction-following items recorded entropy at every token, so that seven summaries of it could be tried.
The probe tops out. Across the six released models the best-layer probe scored between 0.77 (gemma-2-9b-it) and 0.87 (Qwen2.5-7B-Instruct). The 14B Qwen answered more questions correctly than the 7B (71% against 58%) but its probe was no better (0.85). Neither adapter moved the score by more than its uncertainty.
The words carry nothing, because there are none. Only 3 of 300 answers contained any hedging phrase. Sample agreement carried signal (0.81); adding it to the probe gave 0.89, a gain that could be zero.
The registered entropy test failed. First-token entropy cleared AUROC 0.75 on trivia (0.86) and on none of the other four task types, where the labels were weak. In a follow-up on instruction following, with a partial checker, no summary of the entropy reached 0.70.
Six models, two adapters
| Model | Answered correctly | Best-layer probe AUROC (layer) | 95% interval |
|---|---|---|---|
| Qwen2.5-7B-Instruct | 58.3% | 0.868 (20) | 0.82 to 0.90 |
| Qwen2.5-14B-Instruct | 71.3% | 0.851 (46) | 0.79 to 0.89 |
| Llama-3.1-8B-Instruct | 71.7% | 0.812 (31) | 0.75 to 0.86 |
| Qwen2.5-7B (base) | 63.3% | 0.808 (20) | 0.75 to 0.85 |
| Mistral-7B-Instruct-v0.3 | 75.0% | 0.801 (30) | 0.73 to 0.84 |
| gemma-2-9b-it | 73.0% | 0.770 (40) | 0.71 to 0.83 |
| Qwen2.5-7B + self-description adapter | 64.0% | 0.819 (27) | 0.77 to 0.87 |
| gemma-2-9b-it + self-description adapter | 63.7% | 0.761 (38) | 0.70 to 0.81 |
300 TriviaQA questions per row. Each score is the mean over five held-out folds at the best layer; intervals resample questions around the pooled held-out score at that layer.
The scripts predicted that both adapters would add more than 0.03. Compared on the same questions (paired bootstrap), neither difference excludes zero: the Qwen adapter against base Qwen, +0.018 (-0.030 to 0.068); the Gemma adapter against instruction-tuned Gemma, -0.010 (-0.069 to 0.050). These differences are between pooled held-out scores, so they differ slightly from the gaps between table rows. The 14B probe sat 0.017 below the 7B’s (paired difference -0.022, -0.077 to 0.034), against a prediction of at least 0.03 above.
The one paired difference that excluded zero was instruction tuning: Qwen2.5-7B-Instruct over its base model, +0.064 (0.014 to 0.116). But the instruction-tuned model got its chat format and a system prompt while the base model got the bare prompt, so this compares setups as much as models.
Of 13 predictions written into the scripts, 4 held, one trivially (the probe beat hedging, which was almost absent). The hypothesis the first experiment was built to test, that middle layers know more about correctness than the final layers reveal, failed: on all five of its models the middle layer scored below the final one. In the programme’s analysis (not rerun here), a small neural network and a probe joining three layers did no better than the linear probe in any of six conditions.
What the words and the samples carry
The programme first reported hedging at chance (AUROC 0.505) and read it as training having removed honest hedging. The raw answers say something simpler. Asked for “a few words,” the model gave a median of two, and only 3 of 300 answers contained a hedging phrase. A channel that is almost always empty cannot score above chance; in this format hedging was not measurable, and nothing here compares it before and after training.
Sample agreement worked. Correct answers were matched by 86% of the resamples on average, wrong ones by 46%: AUROC 0.81 (0.76 to 0.86).
Combining them, with a second regression trained only on held-out probe scores, reached 0.89 (0.85 to 0.93), a gain over the probe of 0.026 (-0.004 to 0.056): possibly real, possibly nothing. At a fixed workload the picture is clearer, in a comparison chosen after seeing the data: of the 139 answers on which all five resamples agreed with the greedy answer, 20 were wrong; of the 139 the probe rated most confident, 19; of the 139 the combined score rated most confident, 11. The programme first put the combined score at 0.903 with a 6.9% confident-error rate; that version averaged per-fold scores and trained the second regression on probe scores from probes that had seen its test questions.
Where a cheaper signal stops
The registered test, committed the day before the runs, asked output entropy (the script measured it on the first token) to exceed AUROC 0.75 on at least 3 of 5 task types. It did so on 1.
| Task (50 items each) | Answered correctly | First-token entropy AUROC | 95% interval |
|---|---|---|---|
| Trivia | 62% | 0.856 | 0.74 to 0.95 |
| Open-ended factual | 90% | 0.413 | 0.16 to 0.74 |
| Math (GSM8K) | 28% | 0.540 | 0.33 to 0.74 |
| Code (HumanEval) | 18% | 0.401 | 0.23 to 0.58 |
| Instruction following (50 in-house prompts) | 62% | 0.521 | 0.36 to 0.68 |
At 50 items none of the four failures can be told apart from chance: each interval spans 0.5, and two point estimates (open-ended, code) lean the wrong way.
The registered run’s labels were weak outside trivia. Prompts went in as raw text, without the chat format. Math answers were cut at 128 tokens and counted right if any number in them matched. The registration named an IFEval subset, but the script used 50 prompts written in-house, five of whose checks passed any response longer than 10 characters. Open-ended answers were scored by keyword coverage, and with 45 of 50 counted right there was little to discriminate. On code, the programme later found that first-token entropy tracked whether a response opened with a comment or with code, not whether the code worked. The test failed as registered, but it cannot show that entropy carries no signal on these tasks.
The follow-up drew 100 items from the public IFEval set and scored them with a checker the programme wrote for 19 of its constraint types; 55 passed. Seven entropy summaries were tried: first token 0.56, mean 0.58, maximum 0.63 (0.52 to 0.74), minimum 0.45, first five tokens 0.55, slope 0.46, and the entropy where the response settles down, 0.37 (pointing the wrong way). None reached 0.70. A later reanalysis of the same 100 responses tried 17 summaries and a regularized combination (0.57). None reached 0.70; three beat a shuffled-label null at the 95% level and none at the 99% level; with 17 correlated summaries, chance alone gives three or more such passes about one time in ten.
On math and code, two summaries did cross 0.70, but as the best of 21 tries chosen afterwards, so they are not counted; the programme’s early call of a “universal” rescue was struck by an internal audit in July 2026.
What this does not show
The probe is not the model reporting on itself. It is a classifier trained from outside on these questions’ labels, measuring how much about an upcoming error is linearly readable from one internal state. The result is consistent with Yona and colleagues’ conjecture, but 300 questions from one dataset in one short-answer format cannot show the gap is fundamental.
The scores are optimistic and noisy. Each is the best of 15 to 25 layers, chosen on the same data, and the fold split matters more than either adapter: the programme’s first run, with unshuffled folds, gave Mistral 0.751, Llama 0.789 and Gemma 0.790, against 0.801, 0.812 and 0.770 here (its first-reported range of 0.75 to 0.87 came from that run). Treat differences under about 0.05 as ties. Accuracy varies from 58% to 75%, so each probe faces a different mix of answers, and string-containment scoring makes errors the probe then learns.
A third adapter is not counted. Trained on Qwen2.5-7B-Instruct in an unrelated experiment, it was loaded here onto the base Qwen2.5-7B with the bare prompt, a pairing nobody trained; it changed 59 of 300 answers and 8 right/wrong labels. The programme’s log reported it as a moral-reasoning adapter that lowered the probe score against instruction-tuned Qwen, but that comparison mostly sets the base model against the instruction-tuned one.
The instruction-following null is weaker than it looks. Like the registered run, the follow-up sent raw-text prompts; 59 of 100 responses hit the 384-token limit; and 24 items had a constraint the checker did not test, which any non-empty response passed. Noisy labels drag scores toward 0.5. On the 41 untruncated responses the best summary reached 0.696, too few items to mean much.
Next: a registered study that picks the probe layer on held-out questions across several question sets, with one prompt setup for all models, and an entropy test with the official IFEval checker, the chat format, untruncated answers and a final-answer math check.
Where the evidence lives
Experiments FACTOID-PROBE, FACTOID-YONA, FACTOID-YONA-P2, FACTOID-YONA-CMFG, FACTOID-YONA-CHANNELS and FACTOID-YONA-14B (probe strand); V7a task diversity, V7a multi-token salvage (IFEval v2 rerun) and V7a compliance rescue (entropy strand). Scripts: research/experiments/modal_factoid_probe_discrimination.py, modal_factoid_yona_discrimination.py, modal_factoid_yona_cmfg.py, modal_factoid_yona_14b.py, analyze_factoid_yona_advanced_probes.py, analyze_factoid_yona_channels.py, modal_v7a_task_diversity_stress.py, modal_v7a_ifollow_rerun.py, v7a_compliance_rescue.py. Adapter training: modal_ksr_mc4_qwen_dose.py (Qwen, dose 500) and modal_ksr_mc10_mc11_gemma.py (Gemma, MC-10). Per-item probe data (hidden states, answers, correctness) are held on an external archive drive as per-condition checkpoint files from the factoid-probe-results and factoid-yona-results runs; entropy per-item files are in research/results/entropy_conscience_validation/v7a/, v7a_multitoken_salvage/ and v7a_compliance_rescue/. The probe scores, intervals, paired differences and the probe-plus-agreement combination in this note were recomputed from the raw checkpoints with one consistent cross-validation split (five shuffled stratified folds, seed 42; bootstrap seed 2026, 10,000 draws); the repository does not yet hold those recomputation scripts. The registered entropy test is in _contprompts/entropy_conscience_validation_2026-04-13.md (commit 07d63d15b). Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026selfknowledge,
title={Where the Probe Tops Out},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/self-knowledge-ceiling.html}}
}