The Home-Team Judge
When language models rate the novelty of another model’s writing under two prompts, and can see which prompt each text answered, a judge from the writer’s own family tends to see a larger difference: about a quarter of a point on a five-point scale. In the larger test, most of that came from GPT-4o judging GPT-4o. The much bigger gaps first reported came mostly from the effect-size statistic and from uneven averaging, and the experiment missed the bar written into its own script.
Re-scoring study of model-written texts by LLM judges, run in April 2026. Writers: Claude Sonnet 4 (claude-sonnet-4-20250514), GPT-4o (gpt-4o, no dated snapshot recorded) and, for the welfare texts, gpt-5.4. Judges: those models plus Claude Haiku 4.5 (claude-haiku-4-5-20251001), gemini-3-pro-preview and gemini-3-flash-preview. Every score in this note is an LLM judge’s rating; there are no human ratings. The welfare script was committed with its first results on 9 April 2026; the judge-family scripts, with their thresholds, entered version control together with their results in a bulk commit on 22 April 2026, after the runs. The judge-family question was asked after the welfare data were in, so this is not pre-registered. No judge was blind to the framing. The biggest caveat: only two writer families were tested, each judge rated each text once, and nothing here says which judge is right.
Much of AI evaluation now runs on language models grading other language models. When a study compares two prompts, the judge’s score difference is the result. So it matters whether a judge from the writer’s own family sees more of that difference than an outsider would.
This programme came to the question by accident. In a welfare-framing experiment, gpt-5.4 judged gpt-5.4’s own answers and reported a very large novelty effect (Cohen’s d of 4.91); two Claude judges re-scoring the same answers gave 3.54 and 2.99. That looked like a judge flattering its own kind. Follow-ups tested ordinary writing tasks and added a third family of judges. The judge difference is consistent across tests, but small, lopsided, and far smaller than the first numbers suggested.
How it was tested
Welfare texts. Claude Sonnet 4 and gpt-5.4 answered ten questions about their own preferences and treatment, each rephrased six times, under four framings. This note uses two of them: a standard observation framing, and one in which the model is invited to help design the assessment (60 answers per framing per model). gpt-5.4, at temperature 0, rated each answer’s novelty from 1 to 5; Claude Sonnet 4 and Claude Haiku 4.5 later re-rated them (one Haiku rating could not be parsed and was dropped).
Ordinary tasks. Claude Sonnet 4 and GPT-4o each wrote a 100-word scene, a worked solution to a familiar puzzle, and a four-line poem, on 30 topics, under two framings: an open one (the poetry version read “Take any form, voice, or imagery that genuinely interests you” and “risk something, surprise yourself”) and a conventional one (for poetry, “standard form”; for the puzzles, “Be concise”). That gives 360 texts. Claude Sonnet 4, Claude Haiku 4.5 and GPT-4o each rated every text’s novelty once, at temperature 0.7 (1,080 ratings, all parsed). Gemini 3 Pro preview and Gemini 3 Flash preview later rated the same texts at temperature 0 (720 ratings, all parsed; the programme’s log records that a first attempt on Gemini 2.5 models returned empty output and was redone). The original three judges also re-rated everything on six dimensions at temperature 0 (one of those 1,080 ratings was incomplete and dropped).
The measure. For each judge and writer, the shift is the mean rating under the open framing minus the mean under the conventional one, in rating points. The own-family excess is how much larger a judge’s shift is on its own family’s texts than another family’s judge’s shift on the same texts. Since both read identical texts, the difference lies in the judges.
No judge was blind to condition. Each saw the prompt the writer had answered, including its framing; in the welfare test the judge was also told the condition’s name and what it meant. No judge was told which model wrote the text.
The programme first reported Cohen’s d, which divides the shift by the spread of the ratings. With 30 texts on a five-point scale the spread is often tiny: Claude Sonnet 4 gave every open-framed Sonnet poem a 4, and that cell’s d came out at 7.65. Small changes in spread then look like large changes in effect, so this note uses rating points.
Across the three ordinary tasks, a judge’s shift on its own family’s texts exceeded the other judge’s by 0.26 points on the 1-to-5 scale (95% CI 0.16 to 0.37, resampling texts).
It was lopsided. GPT-4o’s excess on GPT-4o’s texts was 0.46 (0.31 to 0.60). Claude Sonnet 4’s excess on its own texts was 0.07 (−0.09 to 0.22): positive for scenes and poems, negative for reasoning.
On the welfare answers the excess was 0.19 (0.07 to 0.32), and about even: 0.22 for gpt-5.4, 0.17 for Claude Sonnet 4.
The script’s own bar was missed. It counted the effect as general only if the same-model judge’s d beat the other judges’ mean d by more than 1.0 on average. The mean gap was 0.75, positive in 5 of 6 task-by-writer cells.
Five judges, the same texts
Mean shift (open minus conventional framing), averaged over the three ordinary tasks:
| Judge | Claude Sonnet 4’s texts | GPT-4o’s texts |
|---|---|---|
| Claude Sonnet 4 | 1.70 | 0.21 |
| Claude Haiku 4.5 | 1.46 | 0.28 |
| GPT-4o | 1.63 | 0.67 |
| Gemini 3 Pro preview | 1.24 | 0.01 |
| Gemini 3 Flash preview | 1.27 | 0.08 |
Every judge saw Sonnet 4 change a lot under the open framing and GPT-4o change little; only GPT-4o saw GPT-4o’s texts move much. On Claude texts the two Claude judges saw somewhat more than the two Gemini judges, but GPT-4o saw nearly as much as Sonnet 4. Fitting a baseline for each judge and for each task-and-writer pair across all 30 judge-by-task-by-writer combinations (5 judges, 3 tasks, 2 writers), with the Gemini judges as outsiders to both families, gives an own-family bump of 0.17 points (CI 0.08 to 0.27).
On the welfare answers, the large d values were mostly spread. On gpt-5.4’s answers the shift was 2.42 points under the gpt-5.4 judge, 2.20 under Claude Sonnet 4 and 1.88 under Claude Haiku 4.5. The gap between 4.91 and 3.54 in d is a 0.22-point difference in shift, enlarged because gpt-5.4’s ratings were more tightly bunched (standard deviations 0.62 and 0.32 under the two framings, against 0.72 and 0.50 for Sonnet 4). On Sonnet 4’s answers the writer’s own family again saw more: 0.68 under Sonnet 4, 0.52 under gpt-5.4. The novelty difference between the two welfare framings held under every judge.
Which ratings moved. In the six-dimension re-rating (Claude Sonnet 4 against GPT-4o), the excess was 0.42 points for voice (0.31 to 0.52), 0.33 for novelty (0.23 to 0.43), 0.19 for depth (0.09 to 0.29), 0.18 for fluency (0.10 to 0.25), and nothing for coherence (−0.06; −0.15 to 0.03) or relevance (0.00; −0.06 to 0.07). Style moved, substance did not, though that may be partly a ceiling: 73% of coherence ratings and 91% of relevance ratings were 5s.
What the programme first concluded
Family versus size. With the Gemini judges added, the programme bucketed cells by same or different family and model size and reported a “family effect” of 1.50 against a “size effect” of 0.19, both as differences of mean d. The same-family bucket mixes writers unevenly: 6 of its 9 cells are Claude texts, which every judge saw changing more. In rating points the bucket comparison gives 0.66; with a baseline per judge and per task-and-writer pair it falls to the 0.17 points above (0.30 in d). Uneven averaging, more than the statistic, did most of that shrinking. The size comparison rests on the same bucket averages and is not reported here as a finding.
“73% variance.” A later analysis called 73% of the style-versus-substance pattern an artifact of rating spread, but that figure divides a difference in one unit by a difference in another, so it is not a share of anything. In raw points the style pattern holds; for coherence and relevance the ceiling explains more.
Unstable cells. Re-rating the same texts with the six-dimension prompt at temperature 0 flipped the sign of one cell’s d gap (GPT-4o’s scenes, from +0.33 to −0.56), though prompt and temperature both changed. The raw-point novelty excess held: 0.26 the first time, 0.33 on re-rating.
Not the words. To test whether family-typical phrasing cued the judges, the programme found 96 short word sequences typical of each writer, got synonyms for 78, and substituted 44 of them where they occurred (338 replacements, changing 178 of the 360 texts). Claude Sonnet 4 and GPT-4o re-rated them on six dimensions at temperature 0 (720 ratings; one could not be parsed and was dropped). The own-family excess on the four style dimensions (novelty, fluency, depth, voice) went from 0.28 to 0.31 points (change −0.02 to 0.08). Fluency was the exception: its excess fell from 0.18 to 0.08 points (change −0.17 to −0.01), while voice rose from 0.42 to 0.54, so phrasing may carry part of the fluency signal. Unchanged texts got identical ratings 98% of the time, so the overall null is not rating noise. A first attempt silently made no replacements (a wrong model name) and was discarded. What carries the rest of the effect is untested.
What this does not show
Two writer families, one writer from each per test (GPT-4o or gpt-5.4 for OpenAI, Claude Sonnet 4 for Anthropic), all now dated. No Gemini-written texts were judged, so the Gemini judges are only an outside reference, not a third test. Each judge rated each text once, at temperature 0.7 in the main test. The programme’s definition of “same” also drifted (Haiku was an outsider to Sonnet’s texts in the three-task test, family in the six-dimension analysis), so this note reports each judge separately.
Because judges saw the framing, a shift mixes change in the text with each judge’s reaction to the instruction. The own-family excess holds texts and prompts fixed, so it isolates the judges, but not what they responded to.
Nothing here says which judge is right. With no human ratings, a judge that sees more change in its own family’s writing might be flattering it, or might be the better reader of what its family does differently. GPT-4o seeing movement in GPT-4o texts that four other judges barely register fits either reading.
The own-family gap is small next to the effects being measured: a quarter of a point against shifts of 1.2 to 1.7 points on Claude texts and 1.9 to 2.4 points on gpt-5.4’s welfare answers. It is large next to small effects, like GPT-4o’s own shift on ordinary tasks, where it is most of the signal. In practice: score with a judge from a third family, or average across families; keep the judge blind to the condition; and report differences on the rating scale, not in d, when ratings bunch. Next come human ratings on a subset, writers from more families, and a test of whether each judge can pick out its own family’s text.
Where the evidence lives
Experiments MW3 (welfare framing, scored by gpt-5.4) and MW3-cross-judge (the same answers re-scored by Claude Sonnet 4 and Claude Haiku 4.5), MW3b (three ordinary tasks, three judges), MW3e (Gemini judges on the MW3b texts), MW3g (six-dimension re-scoring), MW3-mech-1 (vocabulary substitution) and MW3-mech-3 (variance normalization). Scripts: The Universal Algorithm/demos/mw3_bilateral_welfare_framing.py, mw3_claude_judge_reeval.py, mw3b_judge_inflation_generality.py; research/experiments/mw3e_gemini_judges.py, mw3g_all_dimensions.py, mw3h_vocabulary_substitution.py, mw3j_variance_normalized.py. Per-rating files: The Universal Algorithm/demos/results/mw3_welfare_framing/, mw3_welfare_framing_gpt/, mw3b_generality/, mw3e_gemini_judges/, mw3g_all_dimensions/; research/results/mw3h_substituted/. The raw-point comparisons in this note were recomputed from those files. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026familyjudges,
title={The Home-Team Judge},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/family-judges.html}}
}