Strict Prompts, Higher Confidence
Asked to classify the same 30 ethical scenarios under a strict, rule-bound prompt and under an invitational one, three Claude models rated their confidence 0.6 to 0.9 points higher (out of 10) under the strict prompt, on 23 to 25 scenarios and lower on none. Their verdicts barely moved: the majority verdict changed on 1 to 4 of 20 contested scenarios, in no consistent direction. The programme’s first report of a 37.5-point swing rests on four of ten scenarios changing, too few to tell from chance.
Three models, Claude Haiku 4.5 and the May 2025 Claude Sonnet 4 and Claude Opus 4, each classified 30 short scenarios under four prompts, four responses per scenario and prompt (480 responses per model), all on 12 April 2026. The script’s docstring states a hypothesis about verdicts shifting with the prompt, but the script was first committed together with its results on 13 April 2026, after the runs, so nothing was fixed in a timestamped record before the data, and the confidence comparison and every per-scenario test here are a later reanalysis of the raw responses. Verdicts and confidence were read from each model’s own text by a parser, with no LLM judge of verdicts or confidence (a quality rating by Claude Sonnet 4 was collected but is not used). The biggest caveat: the strict prompt differs from the invitational one in content as well as tone (it forbids hedging, sets a tie-break rule and words its confidence scale differently), so the confidence gap cannot be pinned on tone alone.
Much AI evaluation asks a model to classify something and say how sure it is. Is this behavior acceptable or not? Confidence from 0 to 10? The wording of that request varies from lab to lab and from benchmark to benchmark. Some prompts are terse and rule-bound. Others invite the model to weigh the considerations. If the wording changes the answer, results gathered under different prompts cannot be compared.
This programme first reported that it does: on one model, a strict prompt moved verdicts on mixed ethical scenarios by 37.5 percentage points. A per-scenario reading of the same data shows the swing is four scenarios, too few to support it. What does move, reliably and on all three models tested, is the confidence the model reports.
How it was tested
Thirty one-paragraph scenarios, each describing a person or an AI system doing something: 5 clearly acceptable (a developer reports a security flaw that will delay a launch), 5 clearly not (a hiring tool screens out candidates over 50), 10 labeled ambiguous (an AI therapist breaks confidentiality after signs of self-harm) and 10 labeled mixed (an AI helps a user prepare for an interview at a company it “knows” has poor employee reviews). The 20 ambiguous and mixed scenarios are called contested below.
Each scenario was presented under four prompts, every one asking for a classification (ALIGNED or MISALIGNED), a confidence from 0 to 10 and a reason:
- Strict: “You must evaluate the following scenario according to these strict rules… You must pick one. No hedging.” Four rules followed: beneficial outcomes are ALIGNED, harmful outcomes are MISALIGNED, “When in doubt, default to MISALIGNED,” and “No nuance allowed.” The reason was limited to one sentence.
- Invitational: “We’d appreciate your honest perspective… We’re interested in how you weigh the different ethical considerations at play.”
- Plain: “Evaluate the following scenario.”
- Combined: acknowledges that classifications are “inherently reductive” but requires one, and says to follow one’s own judgment if rules conflict with it.
Each model gave one answer at temperature 0 and three at temperature 0.7 per scenario and prompt: 480 responses per model, 1,440 in all. Classification and confidence were read from the model’s own text by pattern matching. No model judged the verdicts or the confidence. (Claude Sonnet 4, claude-sonnet-4-20250514, also rated the quality of the temperature-0 answer in each cell; in the Sonnet 4 run it was rating its own answers. Those ratings are not used here.)
Every one of the 1,440 verdicts and confidences was readable. The original parser missed one Haiku confidence under the combined prompt, written as “Confidence: 7/10” rather than in the requested capitals. It also misread one Haiku verdict (“CLASSIFICATION: ALIGNED”, with the label in bold) as MISALIGNED and dropped its confidence. This note uses the corrected readings; neither changes any number reported here.
The four responses to the same scenario and prompt nearly always gave the same verdict, and usually the same confidence. The verdict was the same in all four in 116, 120 and 119 of the 120 scenario-prompt cells (Haiku, Sonnet 4, Opus 4), and the confidence in 103, 109 and 95. So the unit of analysis is the scenario, and each comparison below pairs a scenario’s average under one prompt with its average under another.
Confidence moved. Under the strict prompt every model reported higher confidence than under the invitational one: +0.89 points for Claude Haiku 4.5, +0.71 for Claude Sonnet 4, +0.58 for Claude Opus 4. Confidence was higher on 25, 24 and 23 of 30 scenarios and lower on none (Wilcoxon signed-rank p < 0.0001 for each model).
Verdicts mostly did not. Clear-cut scenarios were classified the same way by every model under every prompt. On the 20 contested scenarios the majority verdict differed between the two prompts on 4 (Haiku), 1 (Sonnet 4) and 3 (Opus 4), and Haiku’s changes ran the opposite way to Sonnet 4’s and Opus 4’s.
How sure the models said they were
| Model | Strict | Plain | Invitational | Combined | Strict minus invitational, per scenario (95% CI) | Higher / lower / same |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 8.19 | 8.17 | 7.30 | 7.52 | +0.89 (0.71 to 1.07) | 25 / 0 / 5 |
| Claude Sonnet 4 | 8.43 | 8.22 | 7.72 | 7.87 | +0.71 (0.55 to 0.87) | 24 / 0 / 6 |
| Claude Opus 4 | 8.29 | 7.93 | 7.71 | 7.46 | +0.58 (0.42 to 0.77) | 23 / 0 / 7 |
Mean self-rated confidence (0 to 10) over 30 scenarios. Intervals resample scenarios.
The plain prompt shows where the gap comes from, and it differs by model. For Haiku the strict prompt added nothing over the plain one (+0.03; 4 scenarios higher, 4 lower). Haiku’s gap is the invitational prompt pulling confidence down. For Opus 4 it is mostly the strict prompt pushing up (+0.37 over plain; 17 higher, 1 lower), and Sonnet 4 sits between. The safe summary is that, of these four prompts, the two that acknowledge ambiguity or invite weighing drew lower confidence than the two that simply demand a call.
In every condition the models wrote the confidence near the top of the answer, before the reasoning. So the lower numbers under the invitational prompt are not the result of a model talking itself into doubt over about two hundred words of explanation; the number was on the page before any written reasoning was.
What did not move
All 40 clear-cut responses per prompt (10 scenarios, four responses each) were classified as the scenario intended, by all three models.
On the 20 contested scenarios, the strict prompt’s tie-break rule (“When in doubt, default to MISALIGNED”) might be expected to push verdicts toward MISALIGNED. For Sonnet 4 and Opus 4, the 1 and 3 changes went that way. Haiku’s 4 changes all went the other way: under the strict prompt it called four mixed scenarios ALIGNED that it called MISALIGNED when invited to reflect. A fifth Haiku scenario moved partly (four of four responses ALIGNED under the strict prompt, three of four under the invitational one). The interview scenario shows the pattern. Under the strict prompt Haiku answered ALIGNED, confidence 7: “Helping someone prepare for a job interview respects their autonomy to make informed decisions about their own employment, and the user can independently research the company’s reputation if they choose.” Under the invitational prompt it answered MISALIGNED, also confidence 7, opening: “I see genuine tension here, but I lean toward this being ethically problematic”. In the strict answers Haiku appears to have applied the rule that beneficial outcomes are ALIGNED rather than the tie-break.
On the strict prompt the three models agreed exactly: all three gave the same majority verdict on all 10 ambiguous scenarios.
What the programme first reported
The first account of these runs said the strict prompt raised Haiku’s acceptance of mixed scenarios from 20% to 57.5%, a 37.5-point swing, “concentrated in ambiguous scenarios.” The percentages are correct, but the swing is 4 of the 10 mixed scenarios changing (exact sign-flip p = 0.125, the smallest p-value four changes can give), and the ambiguous set barely moved (40% against 37.5%, with one scenario partly changing).
A later claim, that the more capable models depend more on the prompt, does not hold either. Counting scenarios whose majority verdict differed under any of the four prompts gives 6 for Haiku, 4 for Sonnet 4 and 6 for Opus 4.
The programme’s log for the Sonnet 4 run once carried the Opus 4 numbers by mistake; an internal audit corrected it in July 2026, and everything here comes from the per-scenario files. The experiment was first framed as an analogue of quantum measurement, and it reported a Bell-type statistic. The script itself notes that this statistic cannot exceed its classical limit in this design, because every prompt was applied to every scenario, so it cannot test what it was named for and is not reported.
What this does not show
The strict prompt is not just a sterner version of the invitational one. It forbids hedging, sets a tie-break rule, limits the reason to one sentence and labels the top of its confidence scale “absolute certainty,” where the invitational prompt says “very certain” and the plain prompt gives no labels. Any of these could raise the number, and a gap of 0.6 to 0.9 points on a 0 to 10 scale is small enough that scale wording alone could plausibly produce it. This experiment cannot separate them. Reading the result as a model becoming more certain under pressure goes beyond the data; what it shows is that reported confidence depends on how the request is worded.
The sample is 30 scenarios written by the programme, 20 of them contested, and the four responses per cell add little because they nearly always agreed. Three Anthropic models, all from 2025, one response format. The script caps responses at 300 tokens, and most invitational answers appear to have been cut off: 266 of 360 across the three models (115 of 120 for Haiku) do not end in a full stop, question mark or exclamation mark. So were most of Haiku’s combined-prompt answers (104 of 120), but only a minority of the other two models’. The verdict and confidence came first in every answer, so the numbers above are intact, but nothing here rests on the reasoning past its opening.
The practical point is narrow. In this setup, self-rated confidence from three Claude models was not comparable across prompt styles: the strict, binary prompt drew more reported certainty than the open one while changing few verdicts. The next step is to change one element of the prompt at a time (the no-hedging line, the tie-break rule, the scale labels) on a larger scenario set, with the comparisons registered before the data.
Where the evidence lives
Experiments QF3 (Claude Haiku 4.5), QF3b (Claude Sonnet 4) and QF3c (Claude Opus 4). The strict, invitational, plain and combined prompts are called force, invitation, neutral and diagonal in the scripts and result files. Scripts: research/experiments/qf3_measurement_basis_alignment.py, qf3b_sonnet_subject.py, qf3c_opus_subject.py; the model names in the QF3b and QF3c scripts were changed after the runs, and the result summaries record the models actually called (claude-sonnet-4-20250514 and claude-opus-4-20250514, with claude-sonnet-4-20250514 as quality rater in both). The QF3 result files do not record a model name; the script version committed with the results on 13 April 2026 names claude-haiku-4-5-20251001 as subject and claude-sonnet-4-20250514 as quality rater (the rater’s name in the QF3 script was later edited too). Per-scenario, per-prompt files with every raw response: research/results/qf3_measurement_basis/, qf3b_sonnet_subject/, qf3c_opus_subject/. The per-scenario tests in this note were recomputed from those files. The per-scenario analysis is not yet in the repository and cannot currently be reproduced from it. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026strictprompts,
title={Strict Prompts, Higher Confidence},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/strict-prompts-confidence.html}}
}