Chapter 4
How does a whole field change its mind in four months?
In May and June of 2024 the newsletter's Discord recap ran to 1,187,928 words. In all of it, the phrase test-time compute appears zero times. July and August add three more occurrences between them, in another million words.
By December the count for a single month is 77. Nothing about the sampling changed. The field acquired a concept.
This chapter is about those four months: the cleanest example in the archive of a technical community revising, in public and at speed, its model of how progress works.
There was one theory of progress and everybody had it: bigger models, more data, more pre-training compute, with the returns predicted by scaling laws. Every headline in the archive from that summer is a variation on it — a larger open model, a cheaper way to serve a large model, a new dataset. The interesting arguments were about whether the curve was bending.
Then, a week before everything changed, the field got a preview of what it wanted, and believed it.
Reflection 70B, by Matt from IT Department. A two-person team announces a
fine-tune of Llama-3.1-70B using a technique they call reflection tuning — training the model
to emit explicit thinking and reflection steps before answering —
and claims frontier-beating results from a small amount of synthetic data. The issue records
the claims and, in the same breath, the objections: contamination concerns, worse coding
performance, results nobody could reproduce. Within days it had fallen apart.
It is easy to read that as a story about a bad actor. The more useful reading is that the field was primed. The idea that a model could get better by thinking longer before answering was so attractive, and so nearly in the air, that a thin claim about it went straight to the top of the newsletter. Six days later the real thing shipped.
On 12 September, OpenAI released o1. The newsletter's lede that evening was four words:
Test-time reasoning is all you need.
AI News, 2024-09-12
In the community surface, mentions of the o-series go from 1.4 per ten thousand words in August to 20.4 in September — a fifteenfold jump in a month, in half a million words a month that nobody wrote for a newsletter. In announcement space the same month goes from 1.0 to 56.5.
Everything before this bought capability at training time: more parameters, more tokens, more GPU-months, and then a fixed cost to answer each request. o1 moved the purchase to inference. The model writes a long internal chain of reasoning — thousands of tokens the user never sees — before producing an answer, and it gets better the longer you let it run.
The consequence for anyone building on it was immediate, and the newsletter noticed it the same evening, in a parenthesis:
Under the hood, o1 is trained for adding new reasoning tokens — which you pay for, and OpenAI has accordingly extended the output token limit to >30k tokens (incidentally this is also why a number of API parameters from the other models like
AI News, 2024-09-12temperatureandroleand tool calling and streaming, but especiallymax_tokens, is no longer supported).
That parenthesis is the whole engineering story of the next two years. A request stopped having a predictable price or latency, and the parameter you would have used to bound it stopped existing.
The same lede also flagged the chart that mattered, which was not a benchmark table:
You are used to new models showing flattering charts, but there is one of note that you don't see in many model announcements, that is probably the most important chart of all … we now have scaling laws for test time compute, and it looks like they scale loglinearly.
AI News, 2024-09-12
Read the four months as a sequence and you can watch the idea propagate through every layer of the field in order — first the model, then the reception, then the open-weights response, then the productisation, then the benchmark that made it undeniable:
13 Sep — Learnings from o1 AMA
AI News headlines, September–December 2024
18 Sep — o1 destroys Lmsys Arena, Qwen 2.5, Kyutai Moshi release
21 Nov — DeepSeek-R1 claims to beat o1-preview AND will be open sourced
28 Nov — Qwen with Questions: 32B open weights reasoning model nears o1 in GPQA/AIME
6 Dec — $200 ChatGPT Pro and o1-full/pro, with vision, without API, and mixed reviews
21 Dec — o3 solves AIME, GPQA, Codeforces, makes 11 years of progress in ARC-AGI
Two of those deserve a second look. The 21 November headline is the archive being written forwards at its most useful: a Chinese lab announced a preview of an open reasoning model and promised to release the weights, two months before the release that reorganised the field. Anyone reading that day had the information. Almost nobody acted on it.
And the 6 December headline contains the phrase mixed reviews next to a $200/month price tag. The archive is full of moments like this — the correct scepticism and the eventual outcome sitting in the same sentence, with no way at the time to tell which was which.
The figure shows the mechanism in an unusual amount of detail, because the community surface is large enough to resolve months. The o-series line is an event: a spike in September, a second at o3 in December. The reasoning line is not an event — it is a level shift that keeps climbing for five months after the launch that triggered it, peaking in February 2025 at 28.8, nearly seven times its August value.
And test-time compute, the phrase that did not exist, tracks neither. It ratchets: 0, 0, 22, 24, 43, 77 raw mentions from July through December. That is what a concept entering a vocabulary looks like as opposed to a product entering a news cycle.
Until o1, making a model better meant making it bigger or training it longer. Both happen before you ever use it, and both are somebody else's capital expenditure. Test-time compute moves the lever to the other side: the model spends longer on your question, producing intermediate reasoning it mostly does not show you, and the answer improves as a function of how long you let it run.
The reason that mattered is in the lede from the day of the launch, and it is not the benchmark scores:
You are used to new models showing flattering charts, but there is one of note that you don't see in many model announcements, that is probably the most important chart of all. Dr Jim Fan gets it right: we now have scaling laws for test time compute, and it looks like they scale loglinearly.
AI News lede, 2024-09-12
A scaling law is a promise that a knob keeps working. Loglinear means each further multiplication of thinking time buys another fixed increment of accuracy — not forever, but predictably enough to plan around. For anyone building on models, that was a new kind of thing to own: quality had become a dial on the caller's side rather than a property of the weights.
The bill arrived the same day. From the community recap of 12 September, hours after the lede above:
Users express frustration over o1's performance in coding tasks, citing hidden 'thinking tokens' leading to unexpected costs, as noted at $60 per million tokens … many users report preferring alternatives like Sonnet.
AI News, Discord recap, 2024-09-12
Both descriptions are of the same mechanism. The reasoning happens as tokens; the tokens are billed; and because the model decides how many to spend, the caller cannot know the price of a request before making it. Announcement space saw a scaling law. Practice space saw a bill it could not forecast, on the same afternoon, about the same feature.
Reasoning language in announcement space runs 15.05 → 23.47 → 40.30 in the first half of 2025, then 35.84, 17.21 and 14.31 in the last half-year of the corpus — a 2.8× fall from the peak. Practice space has the same shape, gentler and earlier: 18.13 in late 2024, a 19.52 peak, then down to 9.84.
A fall that size usually means one of two things: the idea failed, or the idea won so completely that naming it became unnecessary. It falls in every surface, so it is not an artifact of what the newsletter sampled, and there is no headline anywhere announcing that reasoning was a mistake. So the question is whether the thing survived under other names — and it did, three times over.
A vocabulary that did not exist in the first half of 2024 — reasoning_effort,
thinking budgets, extended thinking, hybrid reasoning, /no_think — goes from
0.00 to a peak of 11.20 in announcement space and then settles at 3.20. That
curve is not decline. It is a feature becoming boring: first nobody has the words, then everybody
is discussing the words, then the words are in an API reference and nobody discusses them at all.
The alignment vocabulary of 2024 — RLHF, DPO, PPO, preference optimisation — runs at 9.33 in early 2024 and 0.43 at the end, a twenty-onefold fall. What replaced it is verification: RLVR, verifiable rewards, verifiers, 0.00 rising to 2.35, and the crossover falls in the second half of 2025. GRPO, the specific algorithm that made reasoning training cheap, spikes to 4.58 in late 2025 and settles at 0.43 — the same became-boring curve, one level down.
You can watch the swap happen in the ledes. On 26 November 2024, the newsletter's opening line is a phrase that did not exist in its vocabulary a year earlier:
Reinforcement Learning with Verifiable Rewards is all you need.
AI News, 2024-11-26
This is a real change in what training a model means. Learning from human preference rankings is expensive, subjective and caps out at the quality of your raters. Learning from problems whose answers can be checked by a program is cheap, objective, and scales to as many problems as you can generate. That swap is the reason reasoning models proliferated as fast as they did.
Distillation is how an expensive capability becomes a cheap one, and the archive explains the mechanism in a Discord exchange from June 2024, eight months before it mattered publicly:
A user asked for effective ways to perform knowledge distillation from Llama70b to Llama8b on an A6000 GPU. Others suggested techniques like using token probabilities, minimizing the delta of cross-entropy, and employing logits from a larger model as the ground truth.
AI News, Discord recap, 2024-06-03
That last phrase is the whole idea. Normally you train a model against the correct answer: a single right token, and everything else wrong. Distillation trains the small model against the large model's full distribution instead — not just "the next word is Paris" but the large model's entire sense of how likely every alternative was. That distribution carries far more information per example than a bare label, which is why a student model can reach a surprising fraction of its teacher's ability on a fraction of the compute.
Two consequences follow, and both run through the rest of this book. Capability flows downhill: once a frontier model exists, a much smaller open one trained on its outputs is a matter of time and modest money rather than of research. And capability leaks, because the teacher's outputs are available to anyone who can call the API — which is why distillation ends the corpus as an accusation rather than a technique, with frontier labs alleging that their outputs were used to train competitors.
Distillation — training a small model on a large one's outputs — is old and was not new in 2024. But it goes from 1.08 to 6.62 in announcement space, and from 0.97 to 9.42 in practice space, ending the corpus as one of the densest technical terms in either surface.
Note the direction. It grows toward practice, and it is higher among people running models on their own hardware than among people announcing things. By the chapter-2 test that makes it the most real thing the reasoning turn produced.
The mechanism is worth spelling out, because it explains the direction. A reasoning model emits its chain of thought as text. Text is copyable. A frontier model's traces are a training set for a small model, and a small model trained on them recovers a surprising fraction of the capability at a fraction of the size. Distillation is the pipe through which frontier reasoning reaches a consumer GPU — which is precisely why the practice surface talks about it most.
The word reasoning peaked and fell by a factor of three. The three things it introduced are all still rising.
| What | Announcement | Practice | Outcome |
|---|---|---|---|
| reasoning (the word) | 6.26 → 14.10 | 6.79 → 9.22 | peaked, then absorbed |
| reasoning as a knob | 0.00 → 3.20 | 0.00 → 0.49 | became an API parameter |
| verifiable rewards | 0.00 → 2.35 | 0.00 → 0.20 | replaced the old objective |
| RLHF / DPO / PPO | 8.04 → 0.85 | 0.97 → 0.10 | displaced |
| distillation | 1.08 → 6.62 | 0.97 → 9.42 | grows toward practice |
If you built anything before September 2024 that assumed a model request has a bounded, predictable cost and latency, the reasoning turn broke that assumption and nothing has restored it. Timeouts sized for a one-second completion now sit in front of something that may think for two minutes. Per-request cost ceilings became per-request cost distributions with a long tail decided by the model, not by you. Every reasoning-model integration in production is, in part, a workaround for that.
Four months, one concept, and a permanent change to the cost model of the thing everybody was building on. The vocabulary that arrived in those months — test-time compute, verifiable rewards, reasoning tokens, thinking budgets — did not exist in a million words of community text in the summer before, and by the following spring you could not read a model release without it.