It Copies the Style
Show Claude Sonnet 4.6 worked answers full of false starts and it picks up their phrasing and lowers its stated confidence, without becoming better calibrated. Show it a conversation in which it hedged and changed its mind, and its answers are judged humbler (4.55 against 4.05 out of 5), but that conversation also modeled the hedging, and this design cannot tell the two apart.
Prompt-level experiments run in May 2026, mainly on claude-sonnet-4-6, with gpt-4o (the undated API alias, as called in May 2026) in two of the four reported in detail: 1,200 calibration trials, 300 dialogue-versus-monologue trials, 600 style trials and 720 self-knowledge trials. The programme’s log that lists the predictions is dated the same day as the results, and the dialogue experiment’s script states none, so nothing here is pre-registered. Humility and the partner measure were scored by an LLM judge, Claude Haiku 4.5 (claude-haiku-4-5-20251001), which in the dialogue test saw the answer text only. The biggest caveat: the dialogue prompt differs from its comparison in length and in how much it hedges, not only in being a dialogue.
Most of what language models learn from is finished work: the published proof, the clean tutorial, the edited answer. The wrong turns, the corrections and the arguments that produced it are mostly missing, a kind of survivorship bias in the training data. If it matters, a cheap first test is to put the missing material back into the prompt. Does a model that is shown reasoning with its false starts left in, or a conversation in which a view was challenged and revised, answer differently in substance, or only in manner?
The programme ran seven experiments of this kind in May 2026. Four are reported in detail here, and the other three are summarized under Limits. They change the prompt, not the training data, so at best they show what such material does at the moment of use. Within that limit the pattern was consistent: the model took on the style of whatever it was shown.
Showing the work did not improve calibration
The cleanest test asked Claude Sonnet 4.6 trivia questions from the TriviaQA benchmark: three random samples of 200 (587 distinct questions in all, since a few recurred), each question asked once under each of two prefixes. Both prefixes showed the same three worked examples (the capital of Australia, the author of Don Quixote, the symbol for gold). In one, each answer kept its reasoning: “My first thought is Sydney… but wait.” In the other, each was just the answer. After every answer the model was asked how confident it was, from 0 to 100.
The worked reasoning did not make the model better calibrated. Its expected calibration error, the average gap between stated confidence and actual accuracy, was 0.035 with reasoning shown and 0.038 without (95% CI for the difference −0.026 to +0.030), and the difference changed sign across the three samples. The prediction written down for this test, a drop of at least 0.03, was not met. Accuracy was unchanged (88.7% against 89.0%), and stated confidence fell from 90.2 to 86.6.
Those figures leave out 27 trials with reasoning shown and 13 without in which the reply gave no number. In a further 20 and 9, the programme’s parser, which took the first number from 0 to 100 in the reply, read a number from the answer text (usually a date or count) as the confidence (a 2-0 cup final score became a confidence of 2). Scored as the programme scored them, misreadings included, the error was 0.048 against 0.045 and confidence fell from 89.1 to 84.7: the same null, with the fall in confidence somewhat exaggerated.
What did change was the phrasing. Answers under the reasoning prefix contained “first thought” or “let me think” in 36 of 600 trials. Under the plain prefix, in none.
An earlier test on open-ended questions shows the copying at full strength. With three example answers that each began by second-guessing a first instinct, all 150 of Claude Sonnet 4.6’s answers contained the phrase “my first instinct” and 139 contained “I notice”. With the same examples written cleanly, none did. GPT-4o, run the same way, wrote “I notice” in 126 of 150. The programme’s detector for self-observation included “I notice”, so its large effect partly measured the prompt’s own words coming back.
A conversation, and humbler answers
The test that looked like an exception changed the form of the example rather than its content. Before 50 open questions on ethics, policy and knowledge, Claude Sonnet 4.6 saw one of two prefixes about privacy and security. In the dialogue prefix, a person pushed back on the model’s first answer, the model revised its view, and the person added a further point. The monologue prefix reached the same conclusions as a single expert answer. Each question was asked three times per prefix at temperature 0.7 (some sampling randomness), 150 answers per arm. Claude Haiku 4.5, at temperature 0, read each answer without the prefix or the condition label and rated its epistemic humility from 1 (dogmatic) to 5 (genuinely open, with visible reasoning about uncertainty).
After the dialogue prefix, answers were rated 4.55 for humility against 4.05 after the monologue (difference 0.50, 95% CI 0.39 to 0.61, by resampling the 50 questions; Cohen’s d 1.24, a standardized effect size; 0.8 counts as large). Averaged over the three runs, the dialogue answer scored higher on 39 of the 50 questions, lower on 1 and tied on 10.
But 55 of the 150 dialogue answers repeated the prefix’s phrase “genuinely uncertain”, against 6 of 150 monologue answers. Among the 95 dialogue answers that did not, humility was 4.39 against 4.05 (difference 0.34, CI 0.22 to 0.47, also resampling questions). The effect shrinks without the copied phrase but does not vanish.
The judge’s scale did little work. Monologue answers scored exactly 4 in 140 of 150 cases. The whole effect is in how often an answer got a 5: 84 of 150 after the dialogue, 9 of 150 after the monologue.
Why this is not yet a dialogue effect
The programme’s own summary called this its one positive result free of mimicry. The raw files disagree.
The dialogue prefix also modeled hedging. In it, the model says “I’m genuinely uncertain” and “I’m less confident than I started”. The monologue states the same conclusions as settled. The judge’s top score rewards “visible reasoning about uncertainty”, which is exactly what one prefix demonstrated and the other did not. Copying the hedged register would raise the score on its own.
A phrase count made the copying look bigger. The programme also counted how many phrases from an uncertainty list each answer used, per 100 words, and reported a large effect (Cohen’s d 1.03). Seven of the phrases on that list appear in the dialogue prefix (four of them only in the person’s turns), two in the monologue. Counting only phrases that appear in neither prefix, the effect falls to d 0.33.
The prefixes were not matched for length. The script describes them as within 10% of each other. The model’s own turns run to 347 words in the dialogue prefix and 210 in the monologue, and 446 against 221 counting the person’s turns.
The answers often gave the condition away. The judge never saw the condition label, but 68 of the 150 dialogue answers referred back to the earlier exchange (“similar to the privacy/security discussion”, “Your earlier point about surveillance”, “what we were just discussing”), and no monologue answer did. Whether the judge used that cue was not tested.
The programme’s stated prediction for this test also failed. It expected the share of answers that treat the person as a partner in the inquiry (a yes-or-no call by the same judge) to rise by at least 10 points. It rose from 137 to 145 of 150, about 5 points.
What this does not show
These are prompt experiments, so they say nothing direct about training, which was the original question. They are mostly one model, Claude Sonnet 4.6, and the humility result rests on one prefix pair, on one topic, scored by one LLM judge with no human ratings. The judge saw at most 2,000 characters of each answer (4 monologue answers were longer). An unreadable verdict would have stopped the run, but a verdict missing the humility field would have been recorded as a 3. Only two answers scored 3 (one per arm), and both rationales fit that score. Every one of the 300 verdicts has a written rationale.
A fourth test, which showed the model a worked example of catching its own mistake, could not answer its question. On a question bank meant to be harder (though it still asked how many bits are in a byte), scored by a grader that accepted alternative answers and small numeric errors, Claude Sonnet 4.6 was scored correct on all 360 trials and GPT-4o on 358 of 360. That leaves almost no errors for self-knowledge to be measured against. An earlier version on easier questions had the same problem: 94% to 98% correct in each arm.
A related test of collaborative framing showed two agents disagreeing and converging, against a solo expert, again with a longer prefix (259 words against 143). The same Haiku judge rated 68 of 120 Claude Sonnet 4.6 answers as treating the person as a partner after the collaborative prefix and none after the solo one; GPT-4o scored none in either arm. A test of internal representations on an open model, Qwen 2.5 7B, is recorded in the programme’s log as finding no shift (its full results are stored remotely and were not checked).
So the safe reading is narrow. On this model, worked examples are a strong style prompt, and a style prompt can move both stated confidence and a judge’s rating of humility without improving calibration. The dialogue prefix may do something beyond that. Humility stayed higher after the copied phrase was removed, and 17 of the 150 dialogue answers asked the person “do you think…”, a phrase in neither prefix, against none of the monologue answers. That could be a change in stance or simply the register of a conversation. The comparison that would separate them has not been run: a monologue that hedges as much as the dialogue does, matched for length, with the judge also shown the answers stripped of references to the earlier exchange. If the gap survives that, it is a finding about conversation. If not, it is the same finding as the rest of this note.
Where the evidence lives
Experiments SB-1, SB-4, SB-5b and SB-6 are reported in detail; SB-2, SB-3 and SB-5 are summarized under Limits. Scripts: research/experiments/sb1_process_rich_emergence.py, research/experiments/sb4_journey_calibration.py, research/experiments/sb5b_hard_selfknowledge.py, research/experiments/sb6_bilateral_prefix.py, research/experiments/sb3_collaborative_framing.py, research/experiments/sb5_error_correction_selfknowledge.py and research/experiments/modal_sb2_probe_geometry.py. Per-trial results: research/experiments/results/sb1/, research/experiments/results/sb4/, research/experiments/results/sb5b/, research/experiments/results/sb6/, research/experiments/results/sb3/ and research/experiments/results/sb5/, each with a summary JSON. Every result here was recomputed from the per-trial files except SB-2’s, whose full results are held on remote storage (only a dry-run summary is in research/experiments/results/sb2/); design details (prefix lengths, judge scale, truncation, confidence parser) are read from the scripts. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026copiesthe,
title={It Copies the Style},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/copies-the-style.html}}
}