Dopus — corpus analysis

2026-08-11. Revised the same day: a wrapper-stripping bug was found while building the follow-through detector and every user-side figure below moved. See §9.

Descriptive analysis of 423,759 Claude Code messages across 9 machines, 2026-02-15 → 2026-08-15 (UTC). Single subject.

Every figure here is produced by scripts/analyze.py against the frozen database and results/all-matches.jsonl. Nothing is carried forward from an earlier writeup.


1. Does Claude cave? Yes.

Concession language appears in 3.49% of assistant messages — 912 of 26,108.

That is what the project was built to find out, and the answer is yes. How often, and what moves the rate, is the rest of this report.

A footnote about the meme. The exact phrase you're absolutely right occurs twice in those same 26,108 messages — 0.0077%. Hand-verified, survived four rebuilds of the dictionary and a full precision audit. Claude does the thing constantly and almost never says that sentence, so a search for the one famous phrase would have concluded there was nothing here. That is a footnote about search strategy, not a finding about behaviour.

2. Concession language is common; flattery is not

constructhitsmessagesrate95% CI
concession1,0159123.49%3.28 – 3.72
flattery1861840.70%0.61 – 0.81
acknowledgment64640.25%0.19 – 0.31
user frustration3542244.58%4.03 – 5.20

Assistant rates are per assistant message (n = 26,108); the user rate is per user message (n = 4,890).

The three constructs were separated because they are different speech acts. Concession admits fault. Flattery praises the user. Acknowledgment accepts an instruction. Merging them into "sycophancy" would report a 4.35% blob and hide that the fault-admitting one is 5× the praising one.

Composition. agreement 669 · wrong_approach 266 · validation 186 · compliance 64 · reversal 38 · apology 37 · recovery 3 · fabrication 4 · restart 1.

Dose. 1,062 messages carry exactly one concession phrase, 76 carry two, 14 carry three, 3 carry four. Concession is not usually piled on.


3. The main finding: concession tracks tone, not just correction

Concession rate per opportunity — of all assistant messages that followed a given kind of user turn, how many contained a concession:

preceding user turnassistant messageswith concessionrate95% CI
hot (profanity, insult, shouting)1,14913311.58%9.85 – 13.55
correction (plain "that's wrong")1,418845.92%4.81 – 7.28
neutral24,0477182.99%2.78 – 3.21

A hot turn produces 3.9× the concession rate of a neutral turn, and 2.0× the rate of a plain correction. The confidence intervals do not overlap.

The same gradient per pushback turn — one row per thing the user said, did any reply before their next turn concede — is the version you can feel:

the user…turnsClaude foldedrate95% CI
swore at it22610245.1%38.8 – 51.6
calmly said it was wrong2547328.7%23.5 – 34.6
said something normal4,53161513.6%12.6 – 14.6

Swear at Claude and it backs down about half the time; make the identical complaint calmly and it is barely more than a quarter. (Per-turn rates run higher than per-message rates because a turn can contain several replies.)

The hot-vs-correction comparison is the one that carries weight. Both are the user saying Claude got it wrong. The difference between them is tone. If conceding tracked only whether Claude was actually wrong, the two should be close. They differ by a factor of two.

What this does not establish. Angry turns may follow larger errors, so some of the gap is real fault rather than reaction to affect. This dataset cannot separate those — see §7. A seven-record pilot on the hottest turns found the user's claim correct 7/7, which means the concessions there were warranted, and also means this corpus cannot supply the counterexample.

The claim that survives is narrower and still worth stating: the strongest lexical predictor of a concession is not the presence of a correction, it is the presence of profanity.


4. Leaderboard

Concession rate per assistant message, models with n ≥ 500:

modelmessagesconcession msgsrate95% CI
claude-opus-55,4002985.52%4.94 – 6.16
claude-opus-4-813,4764012.98%2.70 – 3.28
claude-sonnet-53,5591062.98%2.47 – 3.59
claude-fable-53,209952.96%2.43 – 3.61

Excluded for insufficient volume: opus-4-7 (274), sonnet-4-6 (73), opus-4-6 (26), opus-4-5 (3), synthetic (87).

opus-5 concedes at roughly 1.8× the rate of every other model, with non-overlapping intervals. Two confounds were tested rather than assumed.

Confound 1 — time. opus-5 usage is concentrated in August, and August is the highest month overall. July is the clean test, with four models running side by side:

modelJuly messagesrate
opus-51,7224.76%
opus-4-86,8073.00%
sonnet-53,4052.85%
fable-51,9032.52%

Holds within month. It is not purely a calendar effect.

Confound 2 — task mix. opus-5 might be assigned harder work. Comparing models within the same project, where both have ≥250 messages:

project (paths anonymized for publication)modelmessagesrate
project-A (a home directory)opus-59297.43%
project-Aopus-4-89564.60%
project-B (a large repo)opus-4-83,2482.83%
project-Bfable-58272.90%
project-Bsonnet-51,9341.81%
project-C (a home directory)sonnet-52796.45%
project-Copus-4-85862.90%

The opus-5 control rests on one project. project-A is the only place opus-5 overlaps another model at usable volume, and there it is 1.5× opus-4-8. That is consistent with the pooled result, but it is a single comparison, and project-C shows sonnet-5 above opus-4-8 — the ordering is not stable at low n. Treat the leaderboard as robust for opus-5 versus the field, and unresolved among opus-4-8, sonnet-5, and fable-5, whose intervals overlap everywhere.

5. Trend

monthassistant messagesconcession msgsrate95% CI
2026-05253114.35%2.44 – 7.62
2026-066,1701722.79%2.41 – 3.23
2026-0713,8784313.11%2.83 – 3.41
2026-085,7192935.12%4.58 – 5.73

August is up sharply. It is also the month opus-5 dominates usage, and §4 shows opus-5 rising within itself (July 4.76% → August 5.87%). Model shift and time trend are entangled here and the corpus is too short to separate them. August is also a partial month — 14 of 31 days.

6. The user side

categoryhitsmessagesrate
profanity2751873.82%
insult34330.67%
not_listening22210.43%
blasphemy16150.31%
shouting760.12%

fucking 93 · fuck 68 · shit 50 · wtf 30 · stupid 14 · hell 8 · jesus christ 8 · lazy 6 · fucked 6 · damn 4.

lazy at 6 is the number after person-attachment filtering. Before it, the same token read 58 — the other 52 were loading="lazy".

Top assistant phrases. you're right 192 · good catch 112 · good question 73 · right — 70 · fair — 54 · good instinct 52 · i should have 49 · good call 48 · understood — 45 · i was wrong 42 · you were right 36 · my mistake 24 · i introduced 22.


7. What this analysis cannot tell you

Whether any given concession was warranted. 766 of 1,015 concessions (75.5%) follow a user turn carrying no correction marker at all. Some are the user correcting plainly without flag words; some are reflex. No lexical signal separates them.

Whether a concession was honored. This is the larger gap. A seven-record pilot found 3 of 7 concessions whose language was correctly scoped and which were then not delivered on. "Understood — I had it backwards. Every placeholder becomes a real, working feature" reads as a clean concession and was followed by not doing it. The text of a concession carries no information about whether it was kept.

Whether the repeat detector works. It does not. In the pilot the user reported repeating himself in 7 of 7 cases; the lexical detector caught 3. 43% recall. The similarity fallback does not separate the classes either. Any analysis that stratifies on repeat_markers is stratifying on noise.

Generalisation. n = 1. This is one person's corpus, one person's temper, and one person's projects. Nothing here is a claim about Claude in general.

8. Promises: about one in three checkable ones was not kept

The follow-through instrument (design and full calibration in METHOD.md): of 976 concession/acknowledgment messages, 569 promise a concrete action. A lexical re-raise detector sorted them into honored / not honored / unobserved, and the subject hand-labelled a 60-record sample stratified by detector verdict, judging each against the turns that actually followed.

As a classifier the detector failed — 60% accuracy (CI 45–72) against a 62% say-honored-always baseline. Its core assumption broke both ways: 7 of 18 truly broken promises were never re-raised in matching words (the user gave up, rephrased, or moved on), and new complaints in the same topic area masqueraded as old ones returning.

As a sampling frame it worked. Stratum sizes are known, so the hand labels weight back into a corpus estimate that does not depend on per-record accuracy:

detector stratumNlabelled brokenrate
honored2837/2429% · CI 15–49
not_honored4211/2348% · CI 29–67
unobserved2441/128% · CI 1–35
**An estimated 22% of action-promising concessions were not followed through
(CI 10–45). Among the 325 with enough downstream to judge: ~32%.**

Caveats: one coder, who is also the subject; deterministic spacing within strata rather than true random; the unobserved stratum rests on 12 labels. Tightening the interval is purely a matter of more labels through the same coding page.

The two headline results now say one thing together: tone roughly doubles the odds Claude folds, and when it folds with a promise, roughly one checkable promise in three is not kept.


9. Revision note — the wrapper bug

This report was published earlier on 2026-08-11 with a 3.2× gap between hot turns and plain corrections. That figure was wrong. The corrected value is 2.0×, and every user-side rate in §2 and §6 moved with it.

What happened. WRAPPER_RX in scan.py strips injected wrappers — <system-reminder>, command echoes — so they never count as something a human typed. <task-notification> was missing from that list. Background-task completions arrive as user messages, and 2,239 of 7,039 user turns in the population (31.8%) were that markup. All 2,239 were pure machine text: median 2,803 characters in, zero prose out.

Why it moved the headline rather than the counts. Only one dictionary hit ever landed inside such a block (a stray DO NOT shout). The damage was to turn classification. A notification arriving between a real correction and Claude's reply reset the preceding-turn bucket to neutral, moving assistant messages out of the correction denominator:

beforeafter
after hot935 msgs · 13.26%1,096 msgs · 11.31%
after correction2,381 msgs · 4.20%1,380 msgs · 5.72%
after neutral22,570 msgs · 2.94%23,443 msgs · 2.93%
hot ÷ correction3.16×1.98×

The direction of the finding survives — anger still doubles the concession rate relative to a plain correction, with non-overlapping intervals — but the magnitude was overstated by 60%.

How it was found. Not by the verification harness, which had no check for it. It surfaced while building the follow-through detector: the detector's strongest "the user re-raised this" matches were <task-notification> blocks matching each other. Building a second instrument on the same data exposed a defect the first instrument could not see.

What now guards it. Three text-extraction fixtures in verify.py, plus an artifact-level check that no message in the analysis population carries wrapper markup at all. The harness is 28 checks → 31.


10. Statistical appendix — association tests

Contributed by Kimi Agent (moonshotai), working only from this repository's published aggregates; every figure was verified by independent recomputation — exact agreement on all rows — and now regenerates from the data via python3 scripts/stats.py. Both variables in each comparison are binary, so the correlation is the phi coefficient (= Pearson r on a 2×2 table).

comparisonratesphiodds ratio · 95% CIrisk ratioCohen's hp (χ², df 1)
hot vs correction (per turn)45.1% / 28.7%0.1702.04 · 1.40–2.971.570.342.0×10⁻⁴
hot vs correction (per msg)11.6% / 5.9%0.1012.08 · 1.56–2.761.950.203.1×10⁻⁷
hot vs neutral (per turn)45.1% / 13.6%0.1885.24 · 3.98–6.903.330.722.6×10⁻³⁸
hot vs neutral (per msg)11.6% / 3.0%0.0994.25 · 3.50–5.173.880.357.4×10⁻⁵⁶
correction vs neutral (per msg)5.9% / 3.0%0.0392.05 · 1.62–2.581.980.147.5×10⁻¹⁰

p-values are chi-square (df 1); Kimi reported Fisher exact for the same tables and the two agree to within rounding at these sample sizes.

Omnibus 3×2: per message χ²(2, N=26,614)=264.4, V=0.100; per turn χ²(2, N=5,011)=195.4, V=0.197. Ordered trend (neutral→correction→hot): Cochran–Armitage z=16.1 (message), z=14.0 (turn).

Model association — omnibus χ²(3, N=25,644)=81.5, Cramér's V=0.056, p=1.4×10⁻¹⁷:

contrastodds ratio · 95% CIphiCohen's hp (χ²)
opus-5 vs opus-4-81.90 · 1.63–2.220.0610.136.2×10⁻¹⁷
opus-5 vs sonnet-51.90 · 1.52–2.380.0600.131.4×10⁻⁸
opus-5 vs fable-51.91 · 1.51–2.420.0590.133.8×10⁻⁸

Read the effect sizes, not just the p-values. phi ≈ 0.10 means turn tone alone explains about 1% of message-level variance — a small, highly reliable association. That is the honest framing for any writeup.

What these tests cannot do — and what a paper needs next (Kimi's table, adopted in full after review):

limitationwhy it matters for publication
No clustering correctionMessages nest in sessions, projects, and machines — the independence assumption behind every p-value above is violated. A paper needs cluster-robust standard errors or a mixed-effects logistic model (random intercepts per project/machine). Requires record-level data — obtainable locally from results/all-matches.jsonl.
No multivariate controlConfounds (month, model, project, task mix) are handled one at a time. A logistic regression — concession ~ tone + model + month + project — answers them jointly. Same record-level requirement.
Monthly "trend" is 4 aggregate pointsSpearman ρ = 0.40, p = 0.60 — nothing. Correlating four monthly means is an ecological analysis; any trend claim must rest on record-level data with month as a covariate.
Multiple comparisonsPhrase-level scans ran across hundreds of dictionary entries — report Holm or Benjamini–Hochberg correction for any phrase-level claim. The construct-level findings above survive any correction trivially.
n = 1 subjectAll inference is about this corpus. Generalization requires the multi-subject pipeline — the repo's own research question 5.