Interlude IIIan aside on method
Consolidated, with what each one cost.
Seven findings in this work were published and then withdrawn or reversed. That is the honest number, and it is worth setting out in one place — not as a penance, but because they are not seven unrelated mistakes. They are three mistakes, made three times, three times and once, and each of the three has a different detection method that costs almost nothing to run.
| What I published | What was true | Root cause |
|---|---|---|
| OpenAI's share of headlines fell from 18% to 4% | Templated titles rose from 8% to 68% of issues, so every company falls. Through 2024–25 the templating is the editor's own; in 2026 it is the mirror's, and only 8% of that year's issues were actually sent that way. | the archive was a copy, and the copy was lossy |
| Quiet-day language rose sharply, so the field was slowing | Of 397 recovered issues, 196 are stored under a template and 80 of those were sent under a real headline. Worse, the template is not an absence: on the issues genuinely sent that way it marks days that name more companies, not fewer. | the archive was a copy, and the copy was lossy |
| The editorial layer thinned to nothing — a median of 3 words per issue by 2026 | The mirror had stopped carrying it. Restored from the sent emails the 2026 median is 183 words, close to the 2024 peak of 190. | the archive was a copy, and the copy was lossy |
context rot was a rising concern across the field | A Discord idiom. It disappears entirely once chat logs are excluded from the corpus. | the document stopped being the document |
agentic sat next to retrieval-augmented in 2024, so agents absorbed retrieval | A genre artifact of mixing chat and prose in one embedding. Controlled, the 2024 neighbours are low-code and devika. | the document stopped being the document |
| The effective number of companies discussed rose 4.7×, so the field fragmented | 1.1× inside the Twitter recap and 1.3× inside the Reddit recap. The rest was the sampling frame widening from 7 to 12 subreddits and 384 to 544 accounts. | the document stopped being the document |
| Models from Chinese labs have half the shelf life — 85 days against 175 | At family level they are the longest-lived cohort measured, 398 days against 315. | the unit of observation was wrong |
Three of the seven come from treating a database field as though its contents were stable.
The title column of this archive contains a title in 2023 and, increasingly, a
placeholder afterwards, and the human-written opening underneath it goes the same way.
Eight percent of 2023 issues open with a non-title — not much happened today, a quiet weekend, small news items. By the end of the archive it is 85%, including days carrying frontier model launches and billion-dollar funding rounds.
Any series built on titles therefore has a denominator that is quietly filling with blanks. Every company's headline share falls; I picked the two I expected to fall and wrote a story about them. The same mechanism produced a second finding — that quiet-day language was rising, so the field was slowing down — which is a measurement of the template rather than the field.
What it looks like in the file is this. Here is the entire human-written opening of the last issue in the archive, as stored:
a quiet day.
AI News lede, 2026-08-06, as the public mirror carried it
And here is the opening of the same issue, as it was actually sent, on a day the editor titled AMD buys Taalas:
In The Custom ASIC Thesis we said Taalas was worth paying attention to, and in the Inference Inflection we said everything would go vertical. Our Baseten episode had some skeptical counterpoints against etched LLMs, not just custom ASICs, but clearly Lisa Su disagrees for now.
AI News lede, 2026-08-06, as published
Three words against a paragraph collecting on a call made twice. Nothing in the schema, the
types or the row count distinguishes them, and a series that counts words in the
lede field will read the first one as a field going quiet. One consequence for
anyone checking this: the corpus in the repository behind these pages has since been repaired
from the emails, so the placeholder above is no longer there to look at. The repair is the
commit; the mistake is only in this account of it.
And it is worth noting what the reading still missed even after that correction, because it is the subtler half. Not much happened today is a real editorial signal for most of the corpus, not a blank — and on the issues genuinely sent under it, the front matter names more companies than on the headlined days, not fewer. It marks the absence of a lead story rather than the absence of news. I had corrected the artifact and still misread the value underneath it; the first interlude works through that in full.
Detection: read the values, not the schema. Twenty rows, spread across the range, read by a person. It takes an afternoon and nothing else catches this.
What reading the values still missed. Nothing above is wrong, and it is only half the story. Reading the archive carefully establishes that its title field went blank; it cannot establish why, because the answer is not in the archive. Against 397 issues recovered from the sent emails, the stored and published series track each other through 2024 and 2025 — matching exactly in 2025H1, across 118 issues — the editor really was templating more — and then diverge sharply in 2026, where 64% stored meets 4.6% sent. Four consecutive March 2026 issues the archive files as not much happened today went out as Context Drought and NVIDIA GTC: Jensen goes hard on OpenClaw, Vera CPU, and announces $1T sale. So a real trend and a broken pipeline ran in the same column, in opposite directions, and no amount of careful reading within the archive separates them. Detection at this level costs more than an afternoon: it requires a second copy of the data, obtained by a different route.
Three of the seven come from one cause. The archive's issues are assembled from three kinds of source — chat logs, forum threads and posts — and the proportions invert completely across the corpus. The median issue in early 2024 is almost entirely chat; the median issue at the end contains none at all.
A count taken across a whole issue is therefore a weighted average whose weights are moving, and the weights move faster than most of the things being measured. That produced three separate published claims, each with a different flavour of wrong:
Detection: split the corpus by source and re-run. If the finding only exists in the aggregate, it may be a fact about the mixture rather than about the world. This is the single most productive check in the whole project, and not only because it catches errors — splitting by source to remove a confound is what produced the three-surface comparison that most of this book's surviving results depend on.
One of the seven, and it is the one I would most likely make again, because nothing about it looks like an error.
Measuring how long models stay in the conversation, I used the model tag as the subject.
qwen3.5-235b-a22b is a row; qwen3.6 is a row; qwen3.8-max
is a row. That is a defensible choice — those are genuinely different artifacts with different
weights — and it produced a clean result: models from Chinese labs had a median life of 85 days
against 175 for US frontier labs.
Collapse the version and size suffixes so those three rows become one subject, and the cohort with the shortest lives becomes the longest-lived of any measured: 398 days against 315. The ranking inverts, on identical data, because a lab that ships more named checkpoints of the same underlying thing looks like a lab whose models die young.
Detection: ask what one row is, and whether the thing you actually care about is one row or many. If a decision would be made about the family, measure families. The question is not statistical, and no diagnostic will raise it.
Not one of them was a statistical error. Every count was correct. Every difference was far outside anything sampling noise could produce. Every significance test would have passed, every confidence interval would have been narrow, and every one of those numbers would have been describing the wrong quantity.
All seven live in the step before statistics — the one where you decide what to count, over what population, treating what as a subject. That step has essentially no tooling. There is no linter for it, no test that fails, and no warning in any output. It is checked by reading, or it is not checked.
Every one of these errors produced a clean series with no anomalies in it. That is what makes them dangerous: the output of a broken definition looks exactly like the output of a good one.
There is a selection effect worth naming, because it operates on published work generally and not just on this project.
In every one of the seven cases, the wrong finding was more quotable than the correction. “OpenAI's share of the conversation collapsed” is a headline; “headline share is uninformative because the title field is increasingly a placeholder” is a footnote. “Chinese models have half the shelf life” is a slide; “the ranking depends on whether you count checkpoints or families” is a caveat nobody remembers. “The field fragmented fivefold” travels; “diversity is roughly flat once you hold the source constant” does not.
Errors that produce a clean story are therefore selected for at every stage — they survive review longer, they get repeated more, and they are harder to retract because more people have built on them. Nothing about that is specific to this archive. It is a reason to be most suspicious of your best-sounding result, which is precisely the one you least want to re-examine.
Roughly two months, seven findings, and one genuinely load-bearing claim that had already been written up before it collapsed. Set against that, two of the three fixes produced things worth more than what they destroyed: the source split became the comparison this book is organised around, and the family-level grouping turned a wrong recommendation into a right one.
Read twenty rows, spread across the range, before you count anything. Not
the schema, not the aggregate — the values.
Split by source and re-run. If it only exists in the pooled data, it may be
about the pooling.
State what a subject is and ask whether a decision would be made about that
thing or a different one.
Re-read the same twenty rows on a schedule, because all three of these failures
can arrive later in a series that started out fine.
Be most suspicious of the finding you most want to publish.
None of this is sophisticated, and none of it is new. All of it is cheaper than the two months.