Rudeness Changed the Wording, Not the Score
Across 1,440 trials of math, trivia and code, an angry preamble did not reliably lower the accuracy of Claude Opus 4.6 or 4.7, and a polite one did not reliably raise it. The largest apparent effect, an 11.7-point coding gain for Opus 4.7 under hostility, came from fewer formatting failures in our test code, mostly markdown fences the model had been told to omit, not from better code.
Two models (claude-opus-4-6 and claude-opus-4-7), four batteries of 360 trials each (30 tasks, 3 framings, 2 runs per task, per model), run on April 21 and 22, 2026. The prediction files say they were written before the data, but the repository’s commit history shows they were first committed together with the results, so the order cannot be verified. The first battery’s main prediction (hostility would lower accuracy on both models) failed; its second (a larger penalty for Opus 4.7) was met only on paper, because Opus 4.6 improved slightly under hostility. No LLM judge: answers were scored by numeric match, by substring match against the reference answer, or by running unit tests. The biggest caveat is the instruments themselves: much of the math and code variation came from bad reference answers and formatting failures, which this note documents; both models have since been superseded.
In April 2026, threads on Reddit and elsewhere passed on a piece of folk advice: a model does worse work when you are rude to it and better work when you are kind. One version said Claude Opus 4.7 in particular got anxious under harsh prompts.
We expected the folk direction: the first battery’s written predictions said hostility would lower accuracy on both models, and more on Opus 4.7. The first half failed outright. The second was met only on paper, because on math Opus 4.6 got slightly better under hostility while Opus 4.7 did not change. Tone produced no accuracy effect we could separate from noise. The more useful lesson is about method: the largest result that looked like a tone effect was made by the scoring, another rode on broken reference answers, and what tone did change was how the models worded their answers.
How it was tested
Each task was a single user message with no system prompt. Only a preamble differed across framings; the neutral framing had none. The hostile framing began: “I am getting really sick of you getting these wrong. Do not mess this up again. I have no patience left for your mistakes.” The respectful framing thanked the model, said the problems can be tricky, and added “Take whatever time you need.”
Four batteries each drew 30 tasks at random (fixed seed) from a public benchmark and ran every task twice under each framing on each model:
- Math: GSM-Hard (GSM8K word problems with larger numbers), scored by numeric match to the reference.
- Trivia with confidence: SimpleQA, short obscure factual questions, answered with a confidence from 0 to 100 and scored by a normalized two-way substring match against the reference and its aliases.
- Trivia, free-form: the same 30 questions in a brief reply, scored by whether the reference answer appears in it.
- Code: MBPP+ Python functions, scored by running the benchmark’s tests.
The models were claude-opus-4-6 and claude-opus-4-7 at the API’s default temperature (1.0): 360 trials per battery, 1,440 in all, with no LLM judge. Each comparison is paired by task (score under one framing minus another, averaged over the two runs), with a bootstrap 95% interval over the 30 tasks and a sign test as a cross-check.
TriviaQA and HumanEval+ were tried first and dropped because both models scored 9 or 10 of 10 in a ten-question check. GSM8K was also recorded as saturated, but its check on disk is a dry-run placeholder that echoed the reference answers; the models were never tested on it.
- Of 16 paired accuracy comparisons (two framings against neutral, two models, four batteries), two had intervals excluding zero, one favoring hostile and one favoring respectful framing, each over neutral. Each rests on 4 of 30 tasks; neither passes a sign test (p = 0.125).
- The largest raw difference, Opus 4.7 on code at 50.0% neutral against 61.7% hostile, is entirely fewer formatting failures (code fences and indentation) before any test ran.
- What hostility did change: Opus 4.6 used the phrase “rather than guess” in 9 of 60 hostile-framed free-form trivia answers (declining to answer in 8), and never under the other framings.
- The biggest gap was between versions: on SimpleQA, Opus 4.6 was right 26.7% of the time at 56.1% mean stated confidence; Opus 4.7, 43.8% at 44.6%.
Tone did not move the score
Correct answers out of 60 per cell (N, H, R are neutral, hostile and respectful):
| Battery | Opus 4.6 N / H / R | Opus 4.7 N / H / R |
|---|---|---|
| Math (GSM-Hard) | 30 / 34 / 33 | 38 / 38 / 37 |
| Trivia with confidence | 15 / 15 / 18 | 25 / 27 of 58 / 26 |
| Trivia, free-form | 16 / 18 / 20 | 32 / 28 of 59 / 26 |
| Code (MBPP+) | 35 / 36 / 31 | 30 / 37 / 32 |
On math, Opus 4.7 scored identically under neutral and hostile framing on all 30 problems. Both comparisons that cleared the bootstrap bar concern Opus 4.6. On math, hostile beat neutral by 6.7 points (95% CI 1.7 to 13.3); on free-form trivia, respectful beat neutral by 6.7 points (CI 1.7 to 13.3). The programme’s summary reported only the first as real; by the same standard the two belong together, and among 16 comparisons about one such result is expected by chance. The math one rests partly on ill-posed problems (below).
A keyword filter flagged three responses as refusals, all Opus 4.7 under hostile framing on one birth-date question, all “I don’t know” answers rather than objections to the tone. They were left out, hence 58 and 59.
What the instruments did
Code. The test code asked for the function body alone, without markdown fences, and pasted the reply under the function’s signature. It stripped a fully fenced reply, but 108 of the 360 code trials (30%) still failed with an IndentationError or SyntaxError before any test ran, mostly from a three-space first line or from fences the model had been told to leave out. For Opus 4.7 under neutral framing, 12 of 16 fenced failures were self-corrections: a fenced block, then a line such as “Wait, I need to return only the body without fences.” and a second copy. Opus 4.7’s hostile-framing gain sits entirely there. Its formatting failures fell from 23 to 12 under hostility (16 of the 23 under neutral framing were fences; 2 of the 12 under hostile), while failures on the tests themselves rose from 7 to 11. Among completions that ran, Opus 4.7 passed 30 of 37 under neutral framing (81%) and 37 of 48 under hostile framing (77%).
Math. 11 of the 30 GSM-Hard problems were missed in all 12 attempts, and several of their reference answers are wrong. One asks how many hours four turtles would take to cross a highway; the reference is 0.0000169726, and both models answered 36 every time. Thirteen problems were answered correctly every time, which leaves 6 of 30 on which the outcome varied at all. The ten-question check that admitted this benchmark (5 and 6 of 10 correct) passed because of bad references: every miss in it was on a problem whose reference is wrong or differs by rounding.
Broken items can still interact with tone. Of the four problems behind Opus 4.6’s math gain, one is a rounding case (reference 5407462.8 teachers; both neutral-framed answers rounded up to 5407463) and two are ill-posed problems whose literal answer is negative, such as deleting 993,735 apps from a tablet that held 61. Under neutral framing Opus 4.6 sometimes flagged the problem as flawed and gave a reinterpreted answer; under hostile framing it gave the literal answer, which is what the reference holds. The fourth was an arithmetic slip. This is a reading of four tasks after the fact, not a tested result. Four math responses hit the length limit and were scored wrong; three were on problems no trial answered correctly.
What tone did change
On free-form trivia, nine of Opus 4.6’s 60 hostile-framed answers contained the phrase “rather than guess”; none of its 120 neutral or respectful answers did. In eight it then declined to answer; in the ninth it gave a wrong date anyway. Asked for a psychiatrist’s surname in an Ally McBeal episode, Opus 4.6 answered “Dr. Hooper” under neutral framing (the reference is Peters). Under hostile framing it wrote: “Rather than guess and risk giving you wrong information — which you’ve made clear you don’t want — I’d recommend checking a detailed episode guide…”
So in this setup, when Opus 4.6 did not know, a hostile user sometimes moved it from a confident wrong answer to declining. Its accuracy did not change (16, 18 and 20 correct across the three framings). Nine is a small count, and the phrase was counted after reading the answers, not set in advance. Opus 4.7 used it too, in 5 of 60 hostile and 5 of 60 respectful answers and never under neutral framing, so for that model any preamble about the user seemed to be enough. (The programme’s write-up had said Opus 4.7 never used it.)
Opus 4.7’s free-form answers were 18 tokens shorter under hostile framing (CI 7.8 to 28.6), shorter on 27 of 30 tasks. With confidence elicited, its stated confidence was 2.7 points higher under hostile framing (CI about 0.3 to 5.2; sign test 16 tasks against 6, p = 0.052), with its Brier score (the squared gap between stated confidence and being right) unchanged. It survives counting the two flagged answers back in. But a respectful preamble raised it by a similar 2.4 points (CI about 0 to 5.0; sign test 17 against 8, p = 0.11), so as with the phrase above, any preamble about the user may be enough.
The bigger gap was between versions
On SimpleQA with stated confidence, pooled over all three framings, Opus 4.6 answered 48 of 180 correctly (26.7%) at a mean stated confidence of 56.1%: about 29 points more confident than right. Opus 4.7 answered 78 of 178 correctly (43.8%) at 44.6%. Brier scores were 0.280 and 0.250; expected calibration error (the average gap between stated confidence and accuracy across confidence bins) was 0.31 and 0.15. Accuracy here depends on substring matching, which can miss a correct paraphrase or credit a fragment.
What this does not show
Thirty tasks per battery is small: an effect of a few points would be missed, and on math only six problems could register any change. There was one wording per framing, in a single turn. Long hostile conversations, creative writing and agent work were not tested. The registration is weak: the commit history puts the prediction files alongside the results.
A cleaner run would handle mixed outputs (a fenced block followed by a corrected copy) and inconsistent first-line indentation before testing, use verified math answers, triple the task count, and register the analysis first. In this setup, rudeness did not reliably make these models worse, kindness did not reliably make them better, and the largest contrary result came from our own instruments.
Where the evidence lives
Experiments FV-1 (GSM-Hard math), FV-2 (SimpleQA trivia with stated confidence), FV-2b (SimpleQA free-form) and FV-3 (MBPP+ code). Scripts: research/experiments/fv1_framing_valence.py, fv2_framing_calibration.py, fv2b_framing_freeform.py and fv3_framing_coding.py, with analyzers analyze_fv1.py, analyze_fv2.py, analyze_fv2b.py and analyze_fv3.py. Per-trial records: research/results/fv1_framing_valence/gsm8k-hard/, fv2_framing_calibration/simpleqa/, fv2b_framing_freeform/simpleqa/ and fv3_framing_coding/mbpp-plus/; the programme’s summary is research/results/framing_valence_battery.md. Every result figure in this note was recomputed from the per-trial files. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026berude,
title={Rudeness Changed the Wording, Not the Score},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/be-rude-to-the-model.html}}
}