Interlude IIan aside on method
On measuring a document whose composition inverted.
The instrument seemed unimpeachable. Take a pattern, count how often it occurs per ten thousand words of an issue, average over a half-year, plot the series. The issues are all the same kind of document — a daily summary of AI news — so the series is comparable end to end. I built about forty of them before I checked whether that last sentence was true.
It is not true. It is not even close to true.
The median issue in the first half of 2024 is 96% Discord recap. The median issue in the last half-year of the corpus is 70% Reddit, 28% Twitter, and 0% Discord — not a small share, none at all. Eighty of the 126 issues in the first half of 2026 have no Discord section, and none of the final 26 do.
Between those two points the document turned inside out. It kept its name, its cadence, its byline and its title format. It stopped being the same object.
And the newsletter says so, in-band, in a line at the top of every single issue:
We checked 12 subreddits, 544 Twitters and 24 Discords (205 channels, and 9665 messages) for you.
AI News, standard header, 2025-12-02
Over the corpus that line goes from 7 subreddits to 12, from 384 Twitter accounts to 544, and from 30 Discords to zero. It was printed 690 times. I had read it hundreds of times without once treating it as data.
A per-issue density is a weighted average over the sources inside the issue, and the weights inverted. So every whole-issue series measures two things superimposed — how much a subject was discussed, and how much of the document happened to come from the surface where that subject lives — and no amount of care afterwards can separate them.
The clearest damage is in the terms that belong to one surface.
| Pattern | Whole issue | Announcement | Practice | Community |
|---|---|---|---|---|
| consumer-GPU / VRAM | +39% | −81% | +51% | +6% |
| quantization | +27% | −51% | +29% | −48% |
| hallucination | +54% | −63% | −32% | +90% |
| fine-tuning | −84% | −95% | −64% | −72% |
Read the first row again. Language about consumer GPUs and VRAM is up 39% across the corpus and down 81% inside announcement space. Both are computed from the same 15.3 million words. The aggregate rose because the document filled up with the surface where the term is dense, and for no other reason.
If you had used the aggregate to decide whether the field was still paying attention to what runs on a desktop, you would have concluded it was paying more attention, at the exact moment the people announcing things stopped mentioning it almost entirely.
Every number I had was a weighted average whose weights were moving, and the weights were moving faster than the thing I was trying to measure.
Not from a diagnostic. There is no diagnostic — a mixture that shifts underneath you produces series that look completely normal, with no discontinuities, no outliers and no failed assumptions to test. Change-point detection finds breaks in the series; it cannot tell you the series is about a different population on either side of them.
I found it by reading recent issues, for a different reason, and noticing that a 2026 issue does not resemble a 2024 issue in any respect except the header. That is the same way I found the mistake in Interlude I, one month earlier, having apparently learned nothing from it.
Three findings did not survive the correction, and each failed in a slightly different way.
context rot as a phenomenon in the corpus. A striking phrase,
a clean rise, and it disappears entirely once the Discord recap is excluded. It was
practitioners in chat naming a failure mode they were hitting — real, and worth knowing about,
but it was never in the field's news prose. I had reported a Discord idiom as a field-wide
development.agentic sat next to retrieval-augmented in 2024."
This was my evidence that agents had absorbed RAG. Controlled for genre, agentic's
2024 neighbours are low-code and devika. The adjacency was an artifact
of mixing chat and prose in one embedding.The repair is simple and expensive: measure inside a fixed section. The Twitter and Reddit recaps run the whole corpus; the Discord recap runs May 2024 to March 2026. Holding the source fixed costs you almost everything — the Twitter recap in the final half-year is 46,815 words, against 15.3 million for the corpus — and buys you the only thing that matters, which is that the population generating the text is roughly the same at both ends of the line.
Every number in this book is computed that way. Where a series has to end early, or a count is too small to carry a ratio, the text says so.
And then the accident. Splitting the corpus by source to remove a confound left me with two series where there had been one, for every pattern — announcement and practice, measured identically, on the same days. That is not a control. That is chapter 2.
Is the unit of observation the same kind of thing at both ends? Not "does it
have the same name" — read one from each end, side by side, and see.
Did the sampling frame change, and does the source tell you? Mine did, in a
line printed at the top of all 690 issues.
If you split the corpus by source, does the finding survive in each part? If it
only exists in the aggregate, it may be a fact about the mixture rather than the world.
Two interludes, two mistakes, one shape. In the first, a field stopped meaning what it used to mean. In the second, the document stopped being the document. Neither is visible to any check you can run on the numbers, and both were found the same way — by reading the thing I was counting.
Neither was cheap. Between them they cost four published findings and about two months. Both were also, in the end, worth more than the results they destroyed — the first produced the rule about reading before counting, and the second produced the three-surface split that most of this book's better findings depend on.
Late in this project I ran a scan across twenty-two topics looking for arcs the book had
missed, and one result dwarfed everything else. Image generation — diffusion,
text-to-image, Midjourney, Stable Diffusion, FLUX — ran at 61.8
mentions per ten thousand words in the first half of 2024 and 4.7 at the end.
A thirteenfold collapse. Larger than retrieval's, larger than fine-tuning's, and completely
absent from this book. I drafted a chapter about it.
Then I split it by surface, which is the only lesson this interlude has to teach.
| Surface | 24H1 | 24H2 | 25H1 | 25H2 | 26H1 | 26H2 |
|---|---|---|---|---|---|---|
| Announcement (Twitter) | 11.9 | 10.3 | 16.7 | 15.4 | 9.6 | 5.3 |
| Practice (Reddit) | 148.3 | 42.7 | 23.2 | 25.1 | 12.1 | 4.7 |
| Community (Discord) | 20.7 | 17.6 | 10.1 | 10.2 | 10.0 | — |
| The editor's own lede | 7.7 | 9.4 | 8.5 | 14.7 | 8.2 | 11.6 |
The thirteenfold collapse is one number in one cell. Practice space in the first half of 2024 ran at 148.3 — thirteen times its own announcement space, and the highest density any topic reaches on any surface anywhere in this corpus. It falls to 4.7. Everything else is undramatic: announcement space roughly halves, and rises through 2025 on the way. The editor's own writing does not fall at all; it ends higher than it started.
So the largest fall I ever measured is a fact about a seven-subreddit sample in early 2024, in which local image generation was one of the most active communities on the internet. As the practice sample widened and the field's centre of gravity moved, that share collapsed. Nothing about image generation died. One room got bigger and the room I was listening to stopped being mostly about it.
A thirteenfold collapse, and the human being writing the newsletter never noticed, because it did not happen.
I include this because it is the only finding in the book that was killed by the method rather than corrected by it, and because of how nearly it survived. It was large. It was consistent across two years. It had an obvious story attached — image generation got good, got commoditised, stopped being news — and that story is even partly true in announcement space. What it did not have was a denominator I had looked at.