Written ForwardsChapter 2

Chapter 2

Three surfaces

Why does everything in this field feel like hype?

On 5 February 2026 the newsletter's Twitter recap opens like this:

GPT-5.3-Codex shipped in Codex … framed as advancing frontier coding and professional knowledge in one model. Community reaction highlighted that token efficiency and inference speed may be the most strategically important delta versus prior generations, with one benchmark claim: TerminalBench 2 = 65.4% … Reported efficiency improvements: 2.09× fewer tokens versus GPT-5.2-Codex-xhigh on SWE-Bench-Pro, and together with ~40% speedup implies 2.93× faster at ~+1% score.

AI News, Twitter recap, 2026-02-05

The Reddit recap in the same issue, covering the same day, opens like this:

Anyone here actually using AI fully offline? Running AI models fully offline is feasible with tools like LM Studio, which allows users to select models from Hugging Face based on their hardware capabilities, such as GPU or RAM … While coding workflows may need more powerful setups, consulting tasks can be managed with models like gpt-oss-20b in LM Studio.

AI News, Reddit recap, 2026-02-05

Same publication, same day, same editor. Two populations who are not having the same conversation, and are not close to having the same conversation. One is comparing frontier coding models on token efficiency; the other is asking whether a 20-billion-parameter model on a home machine is good enough to be useful without an internet connection.

Neither is wrong. But if you read only the first one — and the first one is what arrives in your feed, your inbox and your board deck — you will believe things about this field that the second one would have corrected.

Three surfaces

The archive does this every day, structurally, because each issue summarises the same twenty-four hours from three different places and keeps them separate:

Three views of one day, produced by three populations with three different incentives, and kept separate on the page. Nobody designed this as a measuring device. It is one anyway.

There is a fourth layer, and it is different in kind. Above all three recaps sits a short passage written by the editor — the headline claim, the judgement, the argument. It is thin, 124,977 words against 15.4 million, and where the other three sample a population, this one samples a person. That makes it a confound when you are measuring the field and a signal when you are measuring the coverage, and it is worth keeping separate for both reasons.

Kept separate, it turns out to be the best-performing layer in the archive, and it is worth saying so at the start rather than letting it accumulate unremarked. Six times in this book the human paragraph gets somewhere before the three machine layers underneath it do, or flags the exact confound that later had to be corrected for:

None of that makes the lede a measuring instrument; a sample of one person is not a population, and a track record assembled after the fact from the cases a book chose to quote is not a batting average. But it is a standing argument against the reflex this book could otherwise encourage — that the machine layers are the data and the human on top is noise to be controlled for. The human layer is the only place in the corpus where anyone commits to a claim early enough to be wrong in public, which is exactly the property the whole book is about.

The measurement

The method is deliberately dull. Take a pattern — a regular expression for agentic|agents?, say — and count how often it occurs per ten thousand words inside a single named recap section, half-year by half-year. Because the section is fixed, the population writing it is roughly fixed too, and a change in the number is a change in what that population talked about rather than a change in the document around it.

Measuring across the whole issue instead would not work, and the reason is worth one sentence: the mix of sources inside an issue changes enormously over three years, so a whole-issue count partly measures which surface the newsletter happened to be sampling that year. Holding the section fixed removes that. The section is the unit throughout this book.

Run it on the two most promoted ideas of the period and the result is not subtle.

Figure 2 · One idea, two surfacesMentions of agents per 10⁴ words, measured inside the Twitter recap and the Reddit recap. Same days, same corpus, same regular expression.
040801202024H12024H22025H12025H22026H12026H2mentions / 10⁴ wordsagentic (announcement)agentic (practice)

Agents rise 5.5× in announcement space between the first half of 2024 and the first half of 2026. In community space, 3.0×. Among people running models on their own machines, 1.2× — which is to say, essentially not at all.

Reasoning, measured the same way, does not do this: 1.2× in announcement space, 2.1× in community space, 1.4× in practice. Two of those three are at or below the 1.23× that a summarizer swap alone produces, so the honest reading is that this test cannot separate reasoning's trajectory from its own instrument. The staircase is flat and may not be a staircase at all. That is the method declining to confirm a theme, which is the only reason to trust it when it does confirm one.

How much of that staircase is the baseline? More than I would like. The comparison starts at 2024H1 because that is where the corpus's section headings settle, which is a reason unrelated to the answer — but it is not the only defensible start. Run the same measurement from the stable publishing regime that begins on 20 May 2024 and the agent gradient gets steeper: 6.8× announcement against 1.1× practice. Run it from 2024H2 and it vanishes: 3.6× against 3.7×.

The practice-side baseline is doing the work. The Reddit recap in the first half of 2024 is only 30,947 words, and its agent density is three times its 2024H2 value, which suppresses the fold. So the honest statement of this chapter's central finding is narrower than the number suggests: agents rose far more in announcement space than in practice space on two of three baselines, and equally on the third. Of every gradient in this book, only RAG's survives all three. analysis/methods/sensitivity.py runs the comparison.

That descending staircase is what hype looks like when you can measure it. It is not that agents are fake; it is that the further you get from the people with something to announce, the smaller the change becomes. If you only ever read announcement space — and announcement space is what shows up in your feed, your inbox and your board deck — you are reading the largest of three numbers and believing it is the only one. That much is worth knowing even where the exact ratio is not stable, because the direction of the error is always the same.

The part that makes it a tool

If every pattern behaved that way, this would be a complaint about marketing rather than a method. The useful discovery is that some patterns run the other way.

Figure 3 · The hype gradientFold change from 2024H1 to 2026H1 within each surface, log scale. Colour encodes the verdict: red shrinks toward practitioners (narrative), teal grows toward them (under-covered), black is flat (real). Community space is the Discord recap, which ends March 2026 — hence 2026H1 rather than 2026H2 throughout.
0.1×no change10×AnnouncementCommunityPracticeChina blocreasoningagentsquantizationfine-tuningRAGfold change, 2024H1 → 2026H1

The Chinese open-weights bloc — Qwen, DeepSeek, Kimi, GLM, MiniMax — rises 4.6× in announcement space and 9.0× in practice space, with community space higher still at 9.9×. Practitioners were running those models, in volume, before the announcement layer had adjusted to them. Quantization is the same shape in miniature: down 20% in announcement space, up 10% in practice. Both are cases where the coverage was behind the ground.

And some things fall everywhere at once. Retrieval-augmented generation drops by roughly 90% in all three surfaces. Fine-tuning falls hardest in announcement space (to a tenth) and much less in practice (to a third) — which is the signature of something that stopped being newsworthy while remaining a job people do.

So the gradient sorts shifts into three kinds, and the sorting is the whole point:

The test

If a shift shrinks as you approach the people doing the work, the announcement layer is ahead of the practitioner layer. Discount it, and go looking for why.
If it grows as you approach them, practitioners are discussing something the coverage has not caught up with. That is the one worth investigating early.
If it moves about the same everywhere, treat it as a field-wide signal — and triangulate it against something that is not discourse before you act.

None of the three tells you what was deployed. All three tell you where to point the next question, which is a more useful thing than it sounds and a smaller one than it reads.

You can run this without an archive. The surfaces exist for every technical field: there is always a layer where things are announced, a layer where practitioners talk to each other, and a layer where people report what broke. Pick two you can read regularly, and measure the same claim in both. Two partial views you can cross-check beat one comprehensive view you cannot — which is, in one sentence, the methodological argument of this entire book.

Why the surfaces disagree

It is worth being precise about the mechanism, because “hype” is a lazy explanation and the real one is more useful.

Announcement space has a launch schedule. Its population is rewarded for novelty, and its volume is roughly proportional to how much capital is chasing a category. Practice space has a hardware constraint. Its population is rewarded for things that work on a 24GB consumer card tonight, and its volume is proportional to how many people are actually doing the thing. Neither is the truth. Announcement space genuinely leads on things that require a frontier lab to make — reasoning models did arrive, months after they were announced. Practice space genuinely leads on things that require nothing but a download, which is exactly why it saw the Chinese open-weights models first.

The gap between the surfaces is not error. It is the lag between what something costs to announce and what it costs to run.

Go back to 5 February 2026 with that in mind. The Twitter recap is measuring token efficiency on SWE-Bench-Pro because its population ships models and needs a number. The Reddit recap is asking whether a 20-billion-parameter model is good enough offline because its population owns one graphics card and no budget. Neither is reporting on the other. Read both for a month and the gap between them tells you more than either does alone — and it is the only thing in this book you can reproduce without an archive, starting tomorrow, with two browser tabs.