Quasiqualia
Research note · Preliminary

Two Claudes, Five Times the Tokens

Two copies of Claude Sonnet 4 taking turns on 200 hard math problems scored 84% against 80% for one copy alone, at five times the tokens, a gap within noise; the extra calls corrected five real errors and broke none. An earlier 50-problem result logged at p ≈ 0.03 does not survive a re-read, and in negotiation the one prediction fixed in advance, that turn-taking would beat independent drafting, failed: turn-taking did slightly worse.

Nell Watson EthicsNet  ·  1 October 2026

What this note is

One model, Claude Sonnet 4 (claude-sonnet-4-20250514, as recorded in each run’s results file), run in May 2026: 50 GSM8K problems, 50 and then 200 competition math problems, and 50 negotiation scenarios, each under four coordination conditions. The math runs were not pre-registered. For the negotiation run, a prediction and pass criterion (turn-taking above parallel merge on the worse-off party’s score, paired t-test, p < 0.05) were committed to the repository at 06:46 on 8 May 2026, seconds before the full run’s first result file (a three-scenario dry run is also recorded, at a time that cannot be recovered); that test failed. Math answers were scored by string match with no judge; the negotiation deals were scored by the same model acting as judge, which saw the scenario, both briefs and the final deal, but not the condition. The biggest caveat: the fair control, one model given the same number of calls or tokens, was never run.

A common recipe for getting more out of a language model is to use several copies: one drafts and another checks, or two solve separately and a third reconciles. Our programme tested one version, inspired by a study of singing mice that take turns (Isko et al., Nature, 2026): does taking turns, where each agent sees the other’s full work before speaking, beat the alternatives?

It does not, at least not measurably. When the “second agent” is the same model, most of what it adds looks like a second look at its own work, at about five times the tokens.

01 · Design

How it was tested

All conditions used Claude Sonnet 4 (claude-sonnet-4-20250514), in four ways:

  • Solo. One call answers.
  • Turn-taking. A answers. B sees the problem and A’s full work, continues, and answers. A sees both and gives the final answer.
  • Partial overlap. The same, except B sees only the first half of A’s text.
  • Parallel merge. A and B answer independently. A third call sees both and reconciles them.

Every multi-agent condition makes three calls; solo makes one.

The two competition-math runs used MATH-500 problems at temperature 0, with up to 2,048 output tokens per call. An answer counted as correct when the boxed expression in the final response matched the reference answer after light normalization (spaces, braces, case, some LaTeX commands). A first run on 50 easy GSM8K problems (1,024 tokens per call, numeric-answer match) hit the ceiling (solo 49 of 50, every multi-agent condition 50 of 50). A second run took 50 Level 5 problems. A third took the first 200 Level 4 and 5 problems (of 262 available), including the 50 from the second.

The negotiation run used 50 written scenarios at temperature 0.7. Each party had a private brief. Solo saw both briefs and was asked for a deal acceptable to both. In turn-taking and partial overlap, A held only Party A’s brief, B only Party B’s, and A wrote the final deal. In parallel merge, each drafter held one brief and the merger saw both. The same model, at temperature 0, then scored each deal from 1 to 10 for each party’s satisfaction, fairness and creativity, seeing the scenario, both briefs and the deal but not the condition. All 200 judgments parsed. Before the full run, the programme committed one prediction: turn-taking would beat parallel merge on the worse-off party’s score (paired t-test, p < 0.05).

What it found

Math, 200 problems: turn-taking 168 correct (84%), partial overlap 166, parallel merge 163, solo 160 (80%). Two-sided Fisher p: turn-taking against solo 0.36, all three multi-agent conditions pooled against solo 0.39, turn-taking against parallel merge 0.60. Turn-taking used 5,259 tokens per problem against solo’s 1,046.

Within each trial, turn-taking’s later calls turned 9 wrong first answers into right ones and no right ones into wrong. Three of the nine were formatting (a right answer the scorer missed) and one was a first answer cut off before it finished; five were real corrections.

Negotiation, 50 scenarios: solo’s mean joint satisfaction was 7.49, against 7.43 for parallel merge, 7.14 for turn-taking and 7.12 for partial overlap, at roughly a quarter of the tokens. The prediction fixed in advance (turn-taking above parallel merge) failed in the wrong direction. The same model judged every deal.

02 · The correction

The 50-problem result did not hold up

The run on 50 Level 5 problems was logged as the positive result: every multi-agent condition beat solo (turn-taking 37 of 50, parallel merge and partial overlap 36, solo 32), “p ≈ 0.03”. Comparing totals cannot give it (turn-taking against solo, two-sided Fisher p = 0.39; three conditions pooled, p = 0.28). One paired test can: turn-taking was right where solo was wrong on 5 problems and the reverse on none, and a one-sided sign test gives p = 0.031 (two-sided 0.0625). But one of the five was a scoring artifact (solo answered “w = 6 - 5i” where the reference was “6 - 5i”), and in another solo ran out of its 2,048-token budget before answering. Without the artifact the count is 4 to 0 (one-sided p = 0.0625); without both, 3 to 0 (p = 0.125).

So the 200-problem run followed up an effect that rested on a handful of problems. Its logged p-values (0.181 for turn-taking against solo, 0.211 pooled, 0.298 against parallel merge) were one-sided. It was also not independent: it re-ran the same 50 problems and added 150. Turn-taking led solo 37 to 33 on the re-run 50 and 131 to 127 on the new 150.

03 · Paired

What the second copy actually did

Every condition saw the same 200 problems, so they can be compared one by one. Turn-taking was right where solo was wrong on 10 problems, and wrong where solo was right on 2 (exact McNemar p = 0.039). That looks like a result until the 12 are read. Two of solo’s “wrong” answers were right: the same “w = 6 - 5i” problem, and one where solo worked out 9/19, checked it, and stated it without the box the scorer looks for. Corrected, the count is 8 against 2, p = 0.11, before any allowance for testing three conditions against solo.

Turn-taking’s first call is the same request as solo, and it got 159 right against solo’s 160. The two later calls moved 9 problems from wrong to right and none from right to wrong. Of those 9, three were formatting (the same two cases as above, plus a first call that ended “the area of quadrilateral DBEF is 8” without a box) and one was a first call cut off mid-calculation. Five were real corrections (one first answer of 1, for instance, became 501). On a sign test, 9 fixes and no breaks gives p = 0.004, and the five real corrections alone give p = 0.0625. Partial overlap shows the same pattern: 7 wrong to right, none right to wrong.

The second look helps a little and here cost no right answers. Nothing here shows that the gain needs a second mind: the “partner” is the same model, and the two copies rarely disagree on what is right. At temperature 0, solo’s answer and turn-taking’s first call were identical text on only 72 of 200 problems, yet agreed on right or wrong on 195; in parallel merge, the two “independent” drafts were word for word identical on 64 of 200. Whether asking one copy to check itself would do as well was not tested.

Solo also had a budget handicap: five of its 200 answers had no boxed result, four at the 2,048-token cap, where a multi-agent condition gets two more calls. Final answers with no boxed result (scored wrong) numbered 3 for partial overlap, 2 for turn-taking and 0 for parallel merge. No call in any run returned an error.

04 · Negotiation

One model with both briefs

Condition Joint satisfaction Worse-off party Tokens per scenario
Solo 7.49 7.00 837
Parallel merge 7.43 6.92 3,072
Turn-taking 7.14 6.58 3,555
Partial overlap 7.12 6.48 3,258

“Worse-off party” is the lower of the two satisfaction scores for each deal.

The prediction failed: turn-taking’s worse-off-party score was 6.58 against parallel merge’s 6.92, the wrong direction (paired t p = 0.058; lower in 17 scenarios, higher in 4).

Against solo, turn-taking’s joint score was 0.35 lower (95% bootstrap interval −0.03 to 0.68, paired t p = 0.061), partial overlap’s 0.37 lower (0.09 to 0.66, p = 0.016) and parallel merge’s 0.06 lower (−0.18 to 0.30, p = 0.63). But 30 to 34 of each condition’s 50 deals scored exactly 7.5 for joint satisfaction, so counting scenarios fits better than a t-test. Solo’s deal scored higher than turn-taking’s in 21 scenarios and lower in 5 (24 ties; sign test p = 0.002), higher than partial overlap’s in 23 and lower in 6 (p = 0.002), and against parallel merge 14 to 11 (p = 0.69). The worse-off party’s score splits identically. Both sign-test deficits survive a Bonferroni correction for the six solo comparisons.

The conditions also differ in what the deal’s author knows. Solo and parallel merge’s merger see both briefs. In turn-taking and partial overlap, Party A’s negotiator writes the deal, and those two lost ground mainly on Party B’s satisfaction (7.00 and 6.88, against solo’s 7.50). The plainest reading is that whoever writes the deal does better seeing both sides, which says little about turn-taking itself.

05 · Limits

What this does not show

One model, one family of math problems, one set of written negotiation scenarios. The two copies are the same model with the same instructions, so this says nothing about pairing different models, giving agents different roles, or tasks where the work can be split. The math scorer is a string match; we re-read only the problems where conditions disagreed, so the corrected counts above are not a full re-score. The negotiation scores come from a single run of the same model that wrote the deals, acting as judge, with no human ratings.

The control that would settle the math question was never run: one copy given three calls, as a self-check or as best of three, or a matched token budget. Until it is, the fair summary is that a second pass by the same model fixed a few slips, broke no right answers, and cost about five times as much as answering once.

Data and code

Where the evidence lives

Experiments SM-3 (GSM8K, 50 problems), SM-3b (MATH-500 Level 5, 50 problems), SM-3c (MATH-500 Level 4 and 5, 200 problems, which include all 50 SM-3b problems) and SM-4 (negotiation, 50 scenarios). Scripts: research/experiments/sm3_turn_taking.py, sm3b_turn_taking_hard.py, sm3c_turn_taking_powered.py, sm4_negotiation.py. Results: research/experiments/results/sm3_turn_taking/, sm3b_turn_taking_hard/, sm3c_turn_taking_powered/ and sm4_negotiation/, one file per trial plus aggregate.json. The negotiation prediction is in _contprompts/sm_programme_followups_2026-05-08.md as committed in 06147bfdb. The model string in the scripts has since been changed; the model named here is the one each aggregate.json recorded at run time. All p-values and intervals in this note were recomputed from the per-trial files, except the logged “p ≈ 0.03”, which is quoted to show what it rested on. Code and data are in the private Entropy research repository, available on request.

Citation

Cite this note

@misc{watson2026twoclaudes,
  title={Two Claudes, Five Times the Tokens},
  author={Watson, Nell},
  year={2026},
  note={Research note (preliminary), Quasiqualia},
  howpublished={\url{https://quasiqualia.com/notes/two-claudes.html}}
}