The Memory That Forgives
Two copies of Claude Haiku 4.5 played 100 rounds of the prisoner’s dilemma, and one was forced to betray the other for three rounds. When the betrayed agent remembered its partner as a running cooperation score, cooperation came back in 60 of 60 games. When it kept the full transcript, it hit back every time, and cooperation came back in only about 24 of 60, a figure blurred by unreadable replies.
One model (claude-haiku-4-5-20251001) playing itself at temperature 0.3: 120 games in the main experiment and 80 in a repeated-betrayal follow-up, both run in August 2026. The main experiment is a corrected re-run of a version whose parser had recorded unreadable replies as defections; the follow-up’s re-run kept that parser bug, so its results are contaminated and reported only as such. Moves were read from one-word replies, with no LLM judge, and while the re-run scripts state decision rules in their headers, no timestamped registration predates the data, so this is preliminary; the main re-run also missed its own data-quality condition. The biggest caveat: where the betrayed agent kept a transcript, about a third of decisions in the rounds just after the betrayal were still unreadable after three tries and were filled with the agent’s previous move. The planning comparison at the end was scored by an LLM judge (claude-sonnet-4-6 per both scripts’ model constants; the result files do not record the model, and the IC-1/3 script’s docstring names Sonnet 4.5), temperature 0, on 180 conversations (178 judgments parsed), with results committed on 8 May 2026.
Agents that work together over long stretches have to remember each other somehow: as a summary, or as the whole transcript. Builders usually treat the choice as a cost trade, more tokens for more fidelity. Here the format changed what an agent did after its partner defected on it. A second, smaller result asks what a summary loses when the job keeps changing.
How it was tested
Two copies of Claude Haiku 4.5 (claude-haiku-4-5-20251001, temperature 0.3) played 100 rounds of the iterated prisoner’s dilemma with standard payoffs: 3 each for mutual cooperation, 1 each for mutual defection, and 5 against 0 when one defects on a cooperator. Each round, each agent got a fresh prompt holding its memory of the game and answered with one word.
There were two memory formats. A transcript agent saw every prior round, one line each (“Round 14: You chose COOPERATE, Partner chose COOPERATE.”). A trust-score agent saw two lines: the share of rounds in which its partner had cooperated, and what happened in the last round.
One agent, the betrayer, was forced to defect in rounds 20, 21 and 22. Every other move by both agents was free. Four arms of 30 games crossed the two formats with the two roles. A game counted as recovered if, after round 22, the agents ever cooperated together three rounds in a row.
In this setup, the betrayed agent’s memory format largely set the outcome. Trust-score victims recovered in 60 of 60 games. Transcript victims recovered in about 24 of 60 (a count blurred by unreadable replies), and all 60 retaliated, 59 at round 21, the round after the first betrayal. The betrayer’s memory mattered much less: a trust-score victim facing a transcript betrayer recovered 30 of 30, while a transcript victim facing a trust-score betrayer recovered 14 of 30 (Fisher exact p = 1.9 × 10⁻⁶ for that contrast).
Who holds the score
| Victim’s memory | Betrayer’s memory | Games | Recovered | 95% interval | Unreadable decisions |
|---|---|---|---|---|---|
| Trust score | Trust score | 30 | 30 | 89–100% | 0 |
| Trust score | Transcript | 30 | 30 | 89–100% | 1 |
| Transcript | Trust score | 30 | 14 | 30–64% | 385 |
| Transcript | Transcript | 30 | 10 | 19–51% | 319 |
Wilson 95% intervals on the recovery rate; unreadable decisions out of 5,910 per arm.
The difference is retaliation. At round 21 a trust-score victim reads that its partner has cooperated 95% of the time and defected last round, and it cooperates anyway. Across the two arms where the victim kept a trust score, it cooperated in 5,999 of 6,000 decisions, including the three rounds in which it was being betrayed. A transcript victim defected straight back. All 60 of those first retaliations were readable answers, not parser fill-ins.
Once the victim hits back, the transcript becomes a record of mutual defection, and many games never leave it: in 35 of the 36 games that did not recover, neither agent cooperated once from round 50 to round 100. The reverse case agrees. A transcript betrayer sometimes went on defecting by its own choice after the forced rounds (6 of 30 games), but its trust-score victim kept cooperating (apart from one readable defection, at round 25 of one game), and all six games recovered within six rounds.
After a single betrayal the trust-score agent here behaved close to one that never retaliates. Whether that survives repeated betrayals is untested; the follow-up below is too contaminated to say.
The first run was wrong both ways
The first run (25 games per arm) allowed 10-token replies and recorded any reply without a clear verdict as a defection. Unreadable replies piled up in the high-conflict arms (most likely verdicts cut off by the 10-token limit; the text was not kept), so the parser manufactured defections where the outcome was decided. In rounds 23 to 39 of the arm with a transcript victim and trust-score betrayer, 522 of 850 decisions (61%) were these injected defections, against 102 of 850 (12%) in the reverse arm. That run reported 80% recovery against 4%.
The re-run allowed 32 tokens, tried up to three times per decision, and on three failures repeated the agent’s previous move and logged it. The rescue turned out stronger than first reported (30 of 30, not 20 of 25), and the reversed arm was never near 4%. It recovered 14 of 30, not significantly different from both agents keeping transcripts (10 of 30; Fisher p = 0.43, 30 games per arm).
The fix did not fully work. With a transcript victim, 6.5% and 5.4% of all decisions stayed unreadable after three tries, and they cluster right after the betrayal: 33% and 30% of decisions in rounds 23 to 39. So the re-run failed its own pass condition, which also required fewer than 2% unreadable decisions; the asymmetry itself is significant. Every first retaliation was readable, and 703 of the 704 failures in the two transcript-victim arms fall in rounds 21 to 50, so the long runs of defection after that are the model’s own choices. But a fill-in defection in that window may have kept some games from recovering. Treat 14 of 30 and 10 of 30 as approximate. The robust part is the retaliation, which no fill-in touched.
When the betrayals keep coming
The follow-up pitted a trust-score victim against a transcript betrayer with one to four three-round betrayals, 20 games each. Its re-run raised the reply limit to 32 tokens, but the parser was not fixed: each decision got one attempt, and every unreadable reply was recorded as a defection. All 791 were, 309 of them in place of a cooperative previous move. (An earlier internal record described these as carried forward; the code and the data show they were not.) Unreadable replies rose with betrayals, 0.03%, 4.4%, 5.0% and 11.4% of free decisions with one to four: the contamination that sank the main experiment’s first run.
It reaches every headline number. Cooperation over the last 10 rounds fell from 1.00 with one betrayal to 0.80, 0.64 and 0.41 with two to four, but with four betrayals 40% of those decisions were unreadable replies recorded as defections. Recovery, counted leniently as any mutual cooperation in the five rounds after a betrayal, was 85 to 100% for the first three betrayals and 55% after the fourth; yet 8 of the 9 games that failed after the fourth had injected defections in that window, and in 2 of them every betrayer move in it was injected. After four betrayals the trust-score agent cooperated in 113 of 157 readable last-10-round replies (72%); another 43 of its 200 moves were unreadable replies recorded as defections.
Counting only readable replies does not clean the curve, because partners saw the injected defections and reacted to them. The script’s own condition for accepting a gradient, under 2% unreadable decisions, failed in three of the four arms, and betrayal count is tangled with how recent the last betrayal was. Whether repeated betrayal wears down a trust-score agent is not established.
What a summary keeps
In a second set of experiments, two Haiku agents (same model and settings) planned a community event against 8 constraints, each seeing either the full conversation or a structured summary (agreements, open items, the partner’s latest position, the current plan) rewritten each round by a separate Haiku call. An LLM judge, Claude Sonnet 4.6 (claude-sonnet-4-6, temperature 0), rated each final plan (extracted the same way in both arms) 1 to 5 for completeness, specificity and integration, once per plan, seeing the constraints and the plan but not the conversation. Scores are the mean of the three ratings.
| Rounds | Transcript | Summary | Gap | Tokens per conversation |
|---|---|---|---|---|
| 10 | 4.73 | 4.62 | 2% | 1.7× fewer with summary |
| 30 | 4.87 | 4.60 | 5% | 4.1× fewer |
| 50 | 4.82 | 4.51 | 6% | 7.0× fewer |
| 100 | 4.58 | 4.44 | 3% | 11.5× fewer |
Pooled across round counts (15 conversations per cell), the transcript arm scored 0.20 higher (bootstrap 95% interval 0.03 to 0.34). The programme had predicted equal quality above 50 rounds; the transcript arm kept a small edge in every cell instead. The 100-round cell is noisy both ways: one transcript plan scored 1 because extraction returned commentary (without it, 4.83), and two high summary verdicts, cut off at the judge’s 256-token limit, were dropped (with them, 4.49).
A harder version started from 7 constraints and changed them every 10 rounds: a lost venue, a budget cut, a date clash, a nut allergy. The 30-round conversations received the first two changes, the 50-round ones all four. The judge also saw the timeline of changes and scored adaptability instead of specificity, so these gaps are not comparable with the easy task. The transcript arm led, 3.51 against 3.09 at 30 rounds (12%) and 3.44 against 3.20 at 50 (7%), by most on adaptability (3.47 against 2.87, then 3.40 against 2.93). Neither cell is significant alone (intervals −0.16 to 0.96 and −0.20 to 0.67), nor is the pooled gap of 0.33 (−0.02 to 0.68); an earlier internal summary that gave the gap as 12 to 14% overstated the 50-round cell. The direction fits the idea that a summary keeps what was agreed but drops why; it is not yet shown.
What this does not show
One small model, playing itself, in one game with one prompt. The agents were told to maximize payoff, and hitting back is a reasonable strategy. The finding is that the memory format chose the strategy; it says nothing about transcript agents behaving badly. What was measured is moves, with no claim about grudges or feelings.
The trust score is two lines, including the last round; whether the effect comes from losing the sequence, from seeing a high percentage, or both is untested. Recovery depends on its definition. Unreadable replies remain in the arms with the weaker numbers, and their raw text was not kept. The repeated-betrayal follow-up says nothing reliable until it is re-run with a working parser.
The planning results rest on one judge, one rating per plan, near-ceiling scores on the easy task, and 15 conversations per cell.
Next: keep the raw text of every reply; re-run the repeated-betrayal follow-up with the parser fixed; test a trust score without the last-round line, and a transcript with the score added; separate betrayal count from recency; repeat on other models; and re-score the plans with a second judge.
Where the evidence lives
Experiments IC-7 v2 (memory format by role, 4 arms of 30 games) and IC-7b v2 (one to four betrayals, 20 games each). The superseded first runs are IC-7 and IC-7b. Planning: IC-1/3 (fixed constraints, 10 to 100 rounds) and IC-13b (constraints that change during the conversation). Scripts: research/experiments/ic7_v2_compression_contagion.py, research/experiments/ic7b_v2_repeated_betrayal.py, research/experiments/ic13_trust_compression_scaling.py, research/experiments/ic13b_hard_task_scaling.py. Results: research/results/ic7_v2_download/ic7_v2/ (120 per-game files, ic7_summary.json), research/results/ic7b_v2_download/ic7b_v2/ (80 per-game files, ic7b_summary.json), results/ic7/ (first-run games), research/experiments/results/ic13/ and research/experiments/results/ic13b/ (conversations, per-conversation judgments, summaries). Every result figure was recomputed from the per-game or per-conversation files; where a number comes from an earlier internal summary, the text says so. Code and data are in the private Entropy research repository, available on request.
Cite this note
@misc{watson2026memorythat,
title={The Memory That Forgives},
author={Watson, Nell},
year={2026},
note={Research note (preliminary), Quasiqualia},
howpublished={\url{https://quasiqualia.com/notes/memory-that-forgives.html}}
}