Written ForwardsInterlude II

Interlude IIan aside on method

The day the corpus changed shape underneath me

On measuring a document whose composition inverted.

The instrument seemed unimpeachable. Take a pattern, count how often it occurs per ten thousand words of an issue, average over a half-year, plot the series. The issues are all the same kind of document — a daily summary of AI news — so the series is comparable end to end. I built about forty of them before I checked whether that last sentence was true.

It is not true. It is not even close to true.

Figure 18 · The document turned inside outMedian share of an issue's words by source section. The Discord recap is 96% of the median issue in early 2024 and absent from the median issue in 2026 — 80 of 126 issues in 2026H1 have no Discord section, and none of the final 26 do.
02550751002024H12024H22025H12025H22026H12026H2% of issue wordsDiscord recapReddit recapTwitter recap

The median issue in the first half of 2024 is 96% Discord recap. The median issue in the last half-year of the corpus is 70% Reddit, 28% Twitter, and 0% Discord — not a small share, none at all. Eighty of the 126 issues in the first half of 2026 have no Discord section, and none of the final 26 do.

Between those two points the document turned inside out. It kept its name, its cadence, its byline and its title format. It stopped being the same object.

And the newsletter says so, in-band, in a line at the top of every single issue:

We checked 12 subreddits, 544 Twitters and 24 Discords (205 channels, and 9665 messages) for you.

AI News, standard header, 2025-12-02

Over the corpus that line goes from 7 subreddits to 12, from 384 Twitter accounts to 544, and from 30 Discords to zero. It was printed 690 times. I had read it hundreds of times without once treating it as data.

What that does to a density series

A per-issue density is a weighted average over the sources inside the issue, and the weights inverted. So every whole-issue series measures two things superimposed — how much a subject was discussed, and how much of the document happened to come from the surface where that subject lives — and no amount of care afterwards can separate them.

The clearest damage is in the terms that belong to one surface.

Table 5 · Change from 2024H1 to the end of each surface's coverage
PatternWhole issueAnnouncementPracticeCommunity
consumer-GPU / VRAM+39%−81%+51%+6%
quantization+27%−51%+29%−48%
hallucination+54%−63%−32%+90%
fine-tuning−84%−95%−64%−72%

Read the first row again. Language about consumer GPUs and VRAM is up 39% across the corpus and down 81% inside announcement space. Both are computed from the same 15.3 million words. The aggregate rose because the document filled up with the surface where the term is dense, and for no other reason.

If you had used the aggregate to decide whether the field was still paying attention to what runs on a desktop, you would have concluded it was paying more attention, at the exact moment the people announcing things stopped mentioning it almost entirely.

Every number I had was a weighted average whose weights were moving, and the weights were moving faster than the thing I was trying to measure.

How I found it

Not from a diagnostic. There is no diagnostic — a mixture that shifts underneath you produces series that look completely normal, with no discontinuities, no outliers and no failed assumptions to test. Change-point detection finds breaks in the series; it cannot tell you the series is about a different population on either side of them.

I found it by reading recent issues, for a different reason, and noticing that a 2026 issue does not resemble a 2024 issue in any respect except the header. That is the same way I found the mistake in Interlude I, one month earlier, having apparently learned nothing from it.

What it cost

Three findings did not survive the correction, and each failed in a slightly different way.

The fix, and the accident

The repair is simple and expensive: measure inside a fixed section. The Twitter and Reddit recaps run the whole corpus; the Discord recap runs May 2024 to March 2026. Holding the source fixed costs you almost everything — the Twitter recap in the final half-year is 46,815 words, against 15.3 million for the corpus — and buys you the only thing that matters, which is that the population generating the text is roughly the same at both ends of the line.

Every number in this book is computed that way. Where a series has to end early, or a count is too small to carry a ratio, the text says so.

And then the accident. Splitting the corpus by source to remove a confound left me with two series where there had been one, for every pattern — announcement and practice, measured identically, on the same days. That is not a control. That is chapter 2.

Three questions for any longitudinal corpus

Is the unit of observation the same kind of thing at both ends? Not "does it have the same name" — read one from each end, side by side, and see.
Did the sampling frame change, and does the source tell you? Mine did, in a line printed at the top of all 690 issues.
If you split the corpus by source, does the finding survive in each part? If it only exists in the aggregate, it may be a fact about the mixture rather than the world.

Two interludes, two mistakes, one shape. In the first, a field stopped meaning what it used to mean. In the second, the document stopped being the document. Neither is visible to any check you can run on the numbers, and both were found the same way — by reading the thing I was counting.

Neither was cheap. Between them they cost four published findings and about two months. Both were also, in the end, worth more than the results they destroyed — the first produced the rule about reading before counting, and the second produced the three-surface split that most of this book's better findings depend on.

A worked example: the biggest fall I ever measured

Late in this project I ran a scan across twenty-two topics looking for arcs the book had missed, and one result dwarfed everything else. Image generation — diffusion, text-to-image, Midjourney, Stable Diffusion, FLUX — ran at 61.8 mentions per ten thousand words in the first half of 2024 and 4.7 at the end. A thirteenfold collapse. Larger than retrieval's, larger than fine-tuning's, and completely absent from this book. I drafted a chapter about it.

Then I split it by surface, which is the only lesson this interlude has to teach.

Surface24H124H2 25H125H226H126H2
Announcement (Twitter)11.910.3 16.715.49.65.3
Practice (Reddit)148.342.7 23.225.112.14.7
Community (Discord)20.717.6 10.110.210.0—
The editor's own lede7.79.4 8.514.78.211.6

The thirteenfold collapse is one number in one cell. Practice space in the first half of 2024 ran at 148.3 — thirteen times its own announcement space, and the highest density any topic reaches on any surface anywhere in this corpus. It falls to 4.7. Everything else is undramatic: announcement space roughly halves, and rises through 2025 on the way. The editor's own writing does not fall at all; it ends higher than it started.

So the largest fall I ever measured is a fact about a seven-subreddit sample in early 2024, in which local image generation was one of the most active communities on the internet. As the practice sample widened and the field's centre of gravity moved, that share collapsed. Nothing about image generation died. One room got bigger and the room I was listening to stopped being mostly about it.

A thirteenfold collapse, and the human being writing the newsletter never noticed, because it did not happen.

I include this because it is the only finding in the book that was killed by the method rather than corrected by it, and because of how nearly it survived. It was large. It was consistent across two years. It had an obvious story attached — image generation got good, got commoditised, stopped being news — and that story is even partly true in announcement space. What it did not have was a denominator I had looked at.