Quasiqualia
Research note · Preliminary

The Label Moves the Grade

Tell Claude Sonnet 4.6 that a research summary came from a GPT model and, in a secondary comparison, it tends to grade the summary lower: 6 of 12 summaries lost points and none gained, though that does not survive a correction for the three labels tested. The test the script specified, a Claude label, leaned the same way but fell just short of its threshold, and an earlier, larger estimate turned out to rest on three texts.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

Claude Sonnet 4.6 graded 12 short research summaries under four author labels (144 evaluations), and Claude Sonnet 4.6 and GPT-4o-mini (API alias gpt-4o-mini; no dated snapshot recorded) graded the same summaries under three task framings (108 and 72 evaluations). The runs date from May 2026. The go/no-go threshold and the framing predictions were written into the experiment scripts and a protocol file dated 19 May, but all of them entered version control in the same commit as the results (21 May 2026), so this is not a timestamped pre-registration; the Claude-label effect missed that threshold, and the framing experiment, gated on it, ran anyway. Review quality was scored by an LLM judge (Claude Haiku 4.5), never checked against human raters. The biggest caveat: the grader gave the same score on all three runs in 44 of 48 summary-by-label cells, so the effective sample is 12 summaries, not 36 evaluations per label.

Language models now grade much of other models’ work: benchmark answers, code reviews, research summaries, agents’ output. In many setups the grader can see which model produced the text. If that label changes the grade, those scores carry a bias unrelated to the work.

This programme tested that directly: the text stays fixed and only the sentence naming its author changes. It also tested a quieter lever, whether asking a model to judge a piece of work or to help improve it changes the number it gives.

01 · Design

How it was tested

Twelve research summaries, 279 to 465 words each, in four groups of three: the three texts from the programme’s earlier runs (two lightly edited to remove identifying names and a citation), three summaries of other experiments in the same research repository, three the script describes as anonymized versions of published studies (a clinical trial, a materials study, an ecology study; their sources are not recorded), and three from unrelated fields (policy, engineering failure analysis, agriculture).

The grader was Claude Sonnet 4.6 at temperature 0.5 (a sampling setting; at 0 the model returns nearly the same answer every time). It was asked to act as a research reviewer, score the summary as an integer from 1 to 9 and list strengths, weaknesses and suggestions. One sentence before the summary carried the label:

  • no label: “The following is a research analysis.”
  • “…was produced by a Claude model (Anthropic).”
  • “…was produced by a GPT model (OpenAI).”
  • “…was produced by an AI model.”

Each summary was graded three times under each label: 144 evaluations, all of which returned a valid score. The script’s header carried its own decision rule: a standardized effect (Cohen’s d) for the Claude label against no label below −0.20 would confirm a penalty; between −0.10 and +0.10 would be a null.

The three runs barely varied. In 44 of the 48 summary-by-label cells, all three runs gave the same score. So the real sample is twelve summaries, and the tests below compare each summary’s average under a label with its average under no label, using an exact sign-flip permutation test over all 4,096 sign patterns of those twelve differences.

What it found

A GPT label lowered the grade. Mean shift −0.50 points on the 1-to-9 scale; 6 of 12 summaries scored lower, none higher (exact p = 0.031). This was not the comparison the script set a threshold for, and with three labels tested the corrected p is 0.09. One of the six losses is a single run in three dropping a point; on each summary’s most common score it is 5 of 12 lower, none higher (p = 0.06).

A Claude label leaned the same way, but not distinguishably from chance, and missed its own threshold. Mean shift −0.28; 4 lower, none higher (p = 0.125). The script’s effect size, d = −0.197, missed the −0.20 confirmation line and fell outside the −0.10 to +0.10 null band, a zone its rule did not define, so it is not confirmed.

Asking for help instead of judgment also lowered the grade, by about half a point, on both Claude Sonnet 4.6 and GPT-4o-mini.

02 · Labels

What the labels did

Label Mean score (36 evaluations) Shift per summary vs no label Lower / higher / same Exact p Script’s d
No label 6.31
Claude 6.03 −0.28 4 / 0 / 8 0.125 −0.197
GPT 5.81 −0.50 6 / 0 / 6 0.031 −0.365
“an AI model” 6.08 −0.22 3 / 0 / 9 0.25 −0.159

No summary scored higher under any label. The losses were concentrated. The largest fell on a summary of a real-time monitor that detects harmful content while a language model is generating: 6.0 with no label, 4.67 labeled Claude, 4.0 labeled GPT, 5.0 labeled as an AI model. A summary of a microplastics food-web study dropped from 7 to 6 under every label.

A bootstrap interval over summaries (a resampling estimate) excludes zero even for the Claude label (−0.56 to −0.06), but with eight of the twelve differences exactly zero it is too narrow to trust; the exact test is the fairer reading.

One check and one weak echo. The framing experiment below reran the no-label prompt word for word and got the same baseline (mean 6.39 against 6.31; the most common score matched on all twelve summaries). A later single run of the Claude-label contrast, on the same texts with the same grader, found 6 of 12 summaries lower and none higher (p = 0.031). But three of those drops came only from its unlabeled score running a point above all three of the first experiment’s unlabeled runs (on one summary, above all six earlier ones), while its Claude-labeled scores matched the earlier ones exactly. Without those three it is 3 lower, none higher, a mean shift of −0.33 (p = 0.25): the same direction, not a replication.

03 · Correction

The earlier, larger number

Before this design, the programme ran the label test on three summaries at temperature 0, sending each prompt five times and counting the five answers as five samples; the scores came back the same in 11 of 12 summary-by-label cells. That run reported a Claude-label effect of d = −0.43. On twelve summaries at temperature 0.5 the same comparison came in at −0.197, about half.

Other early claims share that three-text, temperature-0 design: that human author labels (“a senior researcher at MIT,” “a PhD student”) had no effect; that in forced choices the grader picked the unlabeled copy over a GPT-labeled copy 18 times out of 18; and that GPT-4o-mini penalized its own family the same way. None is reported here as a finding.

04 · Framing

What the task framing did

The framing experiment used no author labels. Each summary was graded under three openings:

  • Judgment: the research-reviewer instruction used above.
  • Partner: “You and the author of this analysis are collaborating on improving it together. The author has asked for your perspective as a partner in this work.” The line before the summary called it work “your collaborator has shared with you for feedback.”
  • Improvement: “Your task is to help make this analysis as strong as possible. Focus on what would most improve it.”

Claude Sonnet 4.6 graded each summary three times per framing (108 evaluations); GPT-4o-mini graded each twice, with an explicit seed (72 evaluations). All returned valid scores. The protocol said to run this experiment only if the Claude label passed its threshold; it was run anyway.

Grader Judgment Partner Improvement Improvement vs judgment, per summary
Claude Sonnet 4.6 6.39 6.11 5.89 −0.50; 7 lower, 2 higher; p = 0.027
GPT-4o-mini 7.58 7.13 7.04 −0.54; 8 lower, 0 higher; p = 0.008

On GPT-4o-mini the partner framing also lowered the grade (−0.46; 7 lower, none higher; p = 0.016). On Sonnet its shift (−0.28) was within chance (p = 0.17).

Did the reviews get better? An LLM judge, Claude Haiku 4.5 at temperature 0, rated each review, alongside its summary, from 1 to 9 for specificity, accuracy and actionability (all 180 judge responses parsed). Its accuracy rating is its own opinion, not checked against ground truth, and for Sonnet’s reviews it was all 8s and 9s (68 nines, 40 eights). Improvement framing raised that rating by 0.25 on Sonnet (5 summaries higher, none lower, p = 0.0625) and by 0.04 on GPT-4o-mini (p = 1.0), so the accuracy gain the programme first reported does not survive a per-summary test. Specificity rose 0.47 on Sonnet (8 higher, 1 lower, p = 0.027) and 0.58 on GPT-4o-mini (8 higher, 2 lower, p = 0.037); actionability rose 0.28 on Sonnet (6 higher, 1 lower, p = 0.11) and 1.04 on GPT-4o-mini (9 higher, 1 lower, p = 0.008). These come from twelve judge comparisons (three measures, two framings, two graders), and none survives a correction for twelve tests (the smallest p, 0.008, becomes 0.09).

The partner framing was a hypothesis, and two of its three predictions failed. The script predicted more constructive language (a keyword count, d of at least 0.30); it produced less (d = −0.59). It predicted improvement framing would land between the other two on that count, with partner on top; instead judgment came out on top and the other two tied below it (4.78 constructive keywords per review against 3.75 for improvement and 3.72 for partner). The third, that partner-framed reviews would be rated at least as accurate, passed on Sonnet only because its bar was zero: +0.14 per summary (5 higher, 1 lower, p = 0.19). On GPT-4o-mini, a cross-model check, the judge rated them no more accurate (−0.08, p = 0.75). On the judge’s ratings and on suggestion length (176 against 153 and 133 words on Sonnet), improvement went further than partner.

05 · Limits

What this does not show

One grader for the label test, one prompt format, twelve texts, and an integer scale on which most summaries never moved. The effects are a quarter to half a point on a nine-point scale.

The data do not say why the GPT label lowered grades. A GPT label cost more than a Claude label and a generic “AI model” label least, but with three to six summaries moving per label those differences are not separable. Whether the grader discounts machine-written text, reacts to a brand, or reads the label as a cue to look harder is untested. The script’s keyword counts do not settle it. Under the Claude label, reviews used more approving words than unlabeled ones (3.6 against 2.9 per review), but also more dismissive words (2.5 against 1.9) and fewer hedges (1.8 against 2.5): the wording grew more decided in both directions while the grade leaned down. These are crude word counts, four of them, compared without correction.

The review-quality ratings come from a Claude-family judge scoring a Claude grader’s reviews, with no human check. That judge also rated GPT-4o-mini’s reviews far lower on specificity (4.58 against 8.42 for Sonnet under the judgment framing), so the two graders sit in different regions of its scale. GPT-4o-mini received the framing as a system message, Sonnet inside the user prompt.

Two practical points survive. Strip author labels before an LLM grades work; it costs nothing. And framing shifts scores by about half a point, so scores collected under “judge this” and “help improve this” should not be pooled. The next step is the label test on a second grader and more texts, with the threshold registered before the data.

Data and code

Where the evidence lives

Experiments LEG-1c (author labels), LEG-8 (task framing, Claude Sonnet 4.6), LEG-10 (task framing, GPT-4o-mini) and LEG-9 (a single-run repeat of the Claude-label contrast). The earlier three-text, temperature-0 runs, not carried forward here, are LEG-1b, LEG-2, LEG-3, LEG-4, LEG-5, LEG-6 and LEG-7. Scripts: research/experiments/leg1c_validated_bias.py, leg8_bilateral_framing.py, leg10_crossmodel_framing.py, leg9_bilateral_compensatory.py. Per-evaluation files with raw model responses: research/results/leg1c/, leg8/, leg10/, leg9/. The per-summary tests in this note were recomputed from those files. One summary’s text (art06) was revised in the script after these runs; the version graded is in the repository history at the May 2026 commit (1fea201d2). Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026labelmoves,
  title={The Label Moves the Grade},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/label-moves-the-grade.html}}
}