Written ForwardsInterlude III

Interlude IIIan aside on method

The seven things I got wrong

Consolidated, with what each one cost.

Seven findings in this work were published and then withdrawn or reversed. That is the honest number, and it is worth setting out in one place — not as a penance, but because they are not seven unrelated mistakes. They are three mistakes, made three times, three times and once, and each of the three has a different detection method that costs almost nothing to run.

Table 14 · Every finding withdrawn or reversed, and why
What I publishedWhat was trueRoot cause
OpenAI's share of headlines fell from 18% to 4%Templated titles rose from 8% to 68% of issues, so every company falls. Through 2024–25 the templating is the editor's own; in 2026 it is the mirror's, and only 8% of that year's issues were actually sent that way.the archive was a copy, and the copy was lossy
Quiet-day language rose sharply, so the field was slowingOf 397 recovered issues, 196 are stored under a template and 80 of those were sent under a real headline. Worse, the template is not an absence: on the issues genuinely sent that way it marks days that name more companies, not fewer.the archive was a copy, and the copy was lossy
The editorial layer thinned to nothing — a median of 3 words per issue by 2026The mirror had stopped carrying it. Restored from the sent emails the 2026 median is 183 words, close to the 2024 peak of 190.the archive was a copy, and the copy was lossy
context rot was a rising concern across the fieldA Discord idiom. It disappears entirely once chat logs are excluded from the corpus.the document stopped being the document
agentic sat next to retrieval-augmented in 2024, so agents absorbed retrievalA genre artifact of mixing chat and prose in one embedding. Controlled, the 2024 neighbours are low-code and devika.the document stopped being the document
The effective number of companies discussed rose 4.7×, so the field fragmented1.1× inside the Twitter recap and 1.3× inside the Reddit recap. The rest was the sampling frame widening from 7 to 12 subreddits and 384 to 544 accounts.the document stopped being the document
Models from Chinese labs have half the shelf life — 85 days against 175At family level they are the longest-lived cohort measured, 398 days against 315.the unit of observation was wrong

The first kind: the archive was a copy, and the copy was lossy

Three of the seven come from treating a database field as though its contents were stable. The title column of this archive contains a title in 2023 and, increasingly, a placeholder afterwards, and the human-written opening underneath it goes the same way.

Figure 25 · One column, two things happening in itShare of issues whose stored title is a template — “not much happened today”, “a quiet weekend” — rather than a description of the day. Through 2024 and 2025 the rise is the editor's own habit and the stored series tracks what was sent. In 2026 the two part company: of the 133 issues of that year recoverable from the sent emails, only 8.3% went out under a templated subject. A real trend and a lossy mirror, in one column, moving opposite ways.
02550752023202420252026as actually sent: 8%stored titles% of issues

Eight percent of 2023 issues open with a non-title — not much happened today, a quiet weekend, small news items. By the end of the archive it is 85%, including days carrying frontier model launches and billion-dollar funding rounds.

Any series built on titles therefore has a denominator that is quietly filling with blanks. Every company's headline share falls; I picked the two I expected to fall and wrote a story about them. The same mechanism produced a second finding — that quiet-day language was rising, so the field was slowing down — which is a measurement of the template rather than the field.

What it looks like in the file is this. Here is the entire human-written opening of the last issue in the archive, as stored:

a quiet day.

AI News lede, 2026-08-06, as the public mirror carried it

And here is the opening of the same issue, as it was actually sent, on a day the editor titled AMD buys Taalas:

In The Custom ASIC Thesis we said Taalas was worth paying attention to, and in the Inference Inflection we said everything would go vertical. Our Baseten episode had some skeptical counterpoints against etched LLMs, not just custom ASICs, but clearly Lisa Su disagrees for now.

AI News lede, 2026-08-06, as published

Three words against a paragraph collecting on a call made twice. Nothing in the schema, the types or the row count distinguishes them, and a series that counts words in the lede field will read the first one as a field going quiet. One consequence for anyone checking this: the corpus in the repository behind these pages has since been repaired from the emails, so the placeholder above is no longer there to look at. The repair is the commit; the mistake is only in this account of it.

And it is worth noting what the reading still missed even after that correction, because it is the subtler half. Not much happened today is a real editorial signal for most of the corpus, not a blank — and on the issues genuinely sent under it, the front matter names more companies than on the headlined days, not fewer. It marks the absence of a lead story rather than the absence of news. I had corrected the artifact and still misread the value underneath it; the first interlude works through that in full.

Detection: read the values, not the schema. Twenty rows, spread across the range, read by a person. It takes an afternoon and nothing else catches this.

What reading the values still missed. Nothing above is wrong, and it is only half the story. Reading the archive carefully establishes that its title field went blank; it cannot establish why, because the answer is not in the archive. Against 397 issues recovered from the sent emails, the stored and published series track each other through 2024 and 2025 — matching exactly in 2025H1, across 118 issues — the editor really was templating more — and then diverge sharply in 2026, where 64% stored meets 4.6% sent. Four consecutive March 2026 issues the archive files as not much happened today went out as Context Drought and NVIDIA GTC: Jensen goes hard on OpenClaw, Vera CPU, and announces $1T sale. So a real trend and a broken pipeline ran in the same column, in opposite directions, and no amount of careful reading within the archive separates them. Detection at this level costs more than an afternoon: it requires a second copy of the data, obtained by a different route.

The second kind: the document stopped being the document

Three of the seven come from one cause. The archive's issues are assembled from three kinds of source — chat logs, forum threads and posts — and the proportions invert completely across the corpus. The median issue in early 2024 is almost entirely chat; the median issue at the end contains none at all.

A count taken across a whole issue is therefore a weighted average whose weights are moving, and the weights move faster than most of the things being measured. That produced three separate published claims, each with a different flavour of wrong:

Detection: split the corpus by source and re-run. If the finding only exists in the aggregate, it may be a fact about the mixture rather than about the world. This is the single most productive check in the whole project, and not only because it catches errors — splitting by source to remove a confound is what produced the three-surface comparison that most of this book's surviving results depend on.

The third kind: the unit of observation was wrong

One of the seven, and it is the one I would most likely make again, because nothing about it looks like an error.

Measuring how long models stay in the conversation, I used the model tag as the subject. qwen3.5-235b-a22b is a row; qwen3.6 is a row; qwen3.8-max is a row. That is a defensible choice — those are genuinely different artifacts with different weights — and it produced a clean result: models from Chinese labs had a median life of 85 days against 175 for US frontier labs.

Collapse the version and size suffixes so those three rows become one subject, and the cohort with the shortest lives becomes the longest-lived of any measured: 398 days against 315. The ranking inverts, on identical data, because a lab that ships more named checkpoints of the same underlying thing looks like a lab whose models die young.

Detection: ask what one row is, and whether the thing you actually care about is one row or many. If a decision would be made about the family, measure families. The question is not statistical, and no diagnostic will raise it.

What all seven have in common

Not one of them was a statistical error. Every count was correct. Every difference was far outside anything sampling noise could produce. Every significance test would have passed, every confidence interval would have been narrow, and every one of those numbers would have been describing the wrong quantity.

All seven live in the step before statistics — the one where you decide what to count, over what population, treating what as a subject. That step has essentially no tooling. There is no linter for it, no test that fails, and no warning in any output. It is checked by reading, or it is not checked.

Every one of these errors produced a clean series with no anomalies in it. That is what makes them dangerous: the output of a broken definition looks exactly like the output of a good one.

Why the wrong versions survived longer

There is a selection effect worth naming, because it operates on published work generally and not just on this project.

In every one of the seven cases, the wrong finding was more quotable than the correction. “OpenAI's share of the conversation collapsed” is a headline; “headline share is uninformative because the title field is increasingly a placeholder” is a footnote. “Chinese models have half the shelf life” is a slide; “the ranking depends on whether you count checkpoints or families” is a caveat nobody remembers. “The field fragmented fivefold” travels; “diversity is roughly flat once you hold the source constant” does not.

Errors that produce a clean story are therefore selected for at every stage — they survive review longer, they get repeated more, and they are harder to retract because more people have built on them. Nothing about that is specific to this archive. It is a reason to be most suspicious of your best-sounding result, which is precisely the one you least want to re-examine.

What it cost

Roughly two months, seven findings, and one genuinely load-bearing claim that had already been written up before it collapsed. Set against that, two of the three fixes produced things worth more than what they destroyed: the source split became the comparison this book is organised around, and the family-level grouping turned a wrong recommendation into a right one.

The whole method, as a checklist

Read twenty rows, spread across the range, before you count anything. Not the schema, not the aggregate — the values.
Split by source and re-run. If it only exists in the pooled data, it may be about the pooling.
State what a subject is and ask whether a decision would be made about that thing or a different one.
Re-read the same twenty rows on a schedule, because all three of these failures can arrive later in a series that started out fine.
Be most suspicious of the finding you most want to publish.

None of this is sophisticated, and none of it is new. All of it is cheaper than the two months.