Chapter 2
Why does everything in this field feel like hype?
On 5 February 2026 the newsletter's Twitter recap opens like this:
GPT-5.3-Codex shipped in Codex … framed as advancing frontier coding and professional knowledge in one model. Community reaction highlighted that token efficiency and inference speed may be the most strategically important delta versus prior generations, with one benchmark claim: TerminalBench 2 = 65.4% … Reported efficiency improvements: 2.09× fewer tokens versus GPT-5.2-Codex-xhigh on SWE-Bench-Pro, and together with ~40% speedup implies 2.93× faster at ~+1% score.
AI News, Twitter recap, 2026-02-05
The Reddit recap in the same issue, covering the same day, opens like this:
Anyone here actually using AI fully offline? Running AI models fully offline is feasible with tools like LM Studio, which allows users to select models from Hugging Face based on their hardware capabilities, such as GPU or RAM … While coding workflows may need more powerful setups, consulting tasks can be managed with models like
AI News, Reddit recap, 2026-02-05gpt-oss-20bin LM Studio.
Same publication, same day, same editor. Two populations who are not having the same conversation, and are not close to having the same conversation. One is comparing frontier coding models on token efficiency; the other is asking whether a 20-billion-parameter model on a home machine is good enough to be useful without an internet connection.
Neither is wrong. But if you read only the first one — and the first one is what arrives in your feed, your inbox and your board deck — you will believe things about this field that the second one would have corrected.
The archive does this every day, structurally, because each issue summarises the same twenty-four hours from three different places and keeps them separate:
Three views of one day, produced by three populations with three different incentives, and kept separate on the page. Nobody designed this as a measuring device. It is one anyway.
There is a fourth layer, and it is different in kind. Above all three recaps sits a short passage written by the editor — the headline claim, the judgement, the argument. It is thin, 124,977 words against 15.4 million, and where the other three sample a population, this one samples a person. That makes it a confound when you are measuring the field and a signal when you are measuring the coverage, and it is worth keeping separate for both reasons.
Kept separate, it turns out to be the best-performing layer in the archive, and it is worth saying so at the start rather than letting it accumulate unremarked. Six times in this book the human paragraph gets somewhere before the three machine layers underneath it do, or flags the exact confound that later had to be corrected for:
None of that makes the lede a measuring instrument; a sample of one person is not a population, and a track record assembled after the fact from the cases a book chose to quote is not a batting average. But it is a standing argument against the reflex this book could otherwise encourage — that the machine layers are the data and the human on top is noise to be controlled for. The human layer is the only place in the corpus where anyone commits to a claim early enough to be wrong in public, which is exactly the property the whole book is about.
The method is deliberately dull. Take a pattern — a regular expression for
agentic|agents?, say — and count how often it occurs per ten thousand words
inside a single named recap section, half-year by half-year. Because the section is
fixed, the population writing it is roughly fixed too, and a change in the number is a change
in what that population talked about rather than a change in the document around it.
Measuring across the whole issue instead would not work, and the reason is worth one sentence: the mix of sources inside an issue changes enormously over three years, so a whole-issue count partly measures which surface the newsletter happened to be sampling that year. Holding the section fixed removes that. The section is the unit throughout this book.
Run it on the two most promoted ideas of the period and the result is not subtle.
Agents rise 5.5× in announcement space between the first half of 2024 and the first half of 2026. In community space, 3.0×. Among people running models on their own machines, 1.2× — which is to say, essentially not at all.
Reasoning, measured the same way, does not do this: 1.2× in announcement space, 2.1× in community space, 1.4× in practice. Two of those three are at or below the 1.23× that a summarizer swap alone produces, so the honest reading is that this test cannot separate reasoning's trajectory from its own instrument. The staircase is flat and may not be a staircase at all. That is the method declining to confirm a theme, which is the only reason to trust it when it does confirm one.
How much of that staircase is the baseline? More than I would like. The comparison starts at 2024H1 because that is where the corpus's section headings settle, which is a reason unrelated to the answer — but it is not the only defensible start. Run the same measurement from the stable publishing regime that begins on 20 May 2024 and the agent gradient gets steeper: 6.8× announcement against 1.1× practice. Run it from 2024H2 and it vanishes: 3.6× against 3.7×.
The practice-side baseline is doing the work. The Reddit recap in the first half of 2024 is
only 30,947 words, and its agent density is three times its 2024H2 value, which suppresses the
fold. So the honest statement of this chapter's central finding is narrower than the number
suggests: agents rose far more in announcement space than in practice space on two of
three baselines, and equally on the third. Of every gradient in this book, only RAG's
survives all three. analysis/methods/sensitivity.py runs the comparison.
That descending staircase is what hype looks like when you can measure it. It is not that agents are fake; it is that the further you get from the people with something to announce, the smaller the change becomes. If you only ever read announcement space — and announcement space is what shows up in your feed, your inbox and your board deck — you are reading the largest of three numbers and believing it is the only one. That much is worth knowing even where the exact ratio is not stable, because the direction of the error is always the same.
If every pattern behaved that way, this would be a complaint about marketing rather than a method. The useful discovery is that some patterns run the other way.
The Chinese open-weights bloc — Qwen, DeepSeek, Kimi, GLM, MiniMax — rises 4.6× in announcement space and 9.0× in practice space, with community space higher still at 9.9×. Practitioners were running those models, in volume, before the announcement layer had adjusted to them. Quantization is the same shape in miniature: down 20% in announcement space, up 10% in practice. Both are cases where the coverage was behind the ground.
And some things fall everywhere at once. Retrieval-augmented generation drops by roughly 90% in all three surfaces. Fine-tuning falls hardest in announcement space (to a tenth) and much less in practice (to a third) — which is the signature of something that stopped being newsworthy while remaining a job people do.
So the gradient sorts shifts into three kinds, and the sorting is the whole point:
If a shift shrinks as you approach the people doing the work, the announcement
layer is ahead of the practitioner layer. Discount it, and go looking for why.
If it grows as you approach them, practitioners are discussing something the coverage
has not caught up with. That is the one worth investigating early.
If it moves about the same everywhere, treat it as a field-wide signal — and
triangulate it against something that is not discourse before you act.
None of the three tells you what was deployed. All three tell you where to point the next question, which is a more useful thing than it sounds and a smaller one than it reads.
You can run this without an archive. The surfaces exist for every technical field: there is always a layer where things are announced, a layer where practitioners talk to each other, and a layer where people report what broke. Pick two you can read regularly, and measure the same claim in both. Two partial views you can cross-check beat one comprehensive view you cannot — which is, in one sentence, the methodological argument of this entire book.
It is worth being precise about the mechanism, because “hype” is a lazy explanation and the real one is more useful.
Announcement space has a launch schedule. Its population is rewarded for novelty, and its volume is roughly proportional to how much capital is chasing a category. Practice space has a hardware constraint. Its population is rewarded for things that work on a 24GB consumer card tonight, and its volume is proportional to how many people are actually doing the thing. Neither is the truth. Announcement space genuinely leads on things that require a frontier lab to make — reasoning models did arrive, months after they were announced. Practice space genuinely leads on things that require nothing but a download, which is exactly why it saw the Chinese open-weights models first.
The gap between the surfaces is not error. It is the lag between what something costs to announce and what it costs to run.
Go back to 5 February 2026 with that in mind. The Twitter recap is measuring token efficiency on SWE-Bench-Pro because its population ships models and needs a number. The Reddit recap is asking whether a 20-billion-parameter model is good enough offline because its population owns one graphics card and no budget. Neither is reporting on the other. Read both for a month and the gap between them tells you more than either does alone — and it is the only thing in this book you can reproduce without an archive, starting tomorrow, with two browser tabs.