Written ForwardsChapter 5

Chapter 5

Seven days in January

What does it look like when something actually breaks through?

Through the first three weeks of January 2025, the newsletter's Discord recap mentions DeepSeek about thirteen times per ten thousand words — the ordinary background rate for a lab that had shipped a well-regarded open model the previous month. On Monday 20 January the same measurement reads 219.9, from 643 mentions in a single day's recap.

Seventeen times the baseline, overnight, in text nobody wrote for a newsletter. That is the largest single-day movement anywhere in this archive.

What follows is the only week in three years where you can watch the entire field reorganise in real time, day by day. It is worth reading closely, because the interesting question is not what happened — everyone knows what happened — but what distinguishes this from the dozens of launches that also spiked and then reverted. The archive answers that, and the answer is not the benchmark scores.

The week

Monday 20 January

DeepSeek R1: o1-level open weights model and a simple recipe for upgrading 1.5B models to Sonnet/4o level. Community density: 219.9.

The lede that evening ran to 541 words, which is unusual, and it is worth reading a piece of it because it is the clearest statement anywhere in the archive of why this particular release was different from the dozens around it:

GRPO is all you need.

DeepSeek actually dropped 8 R1 models — 2 “full” models, and 6 distillations on open models … Surprisingly, MIT licensed rather than custom licenses, including explicit OK for finetuning and distillation.

Pricing (per million tokens): 14 cents input (cache hit), 55 cents input (cache miss), and 219 cents output. This compares to o1 at 750 cents input (cache hit), 1500 cents input (cache miss), 6000 cents output. That's 27x–50x cheaper than o1.

R1 distillations were remarkably effective, giving us this insane quote: “DeepSeek-R1-Distill-Qwen-1.5B outperforms GPT-4o and Claude-3.5-Sonnet on math benchmarks with 28.9% on AIME and 83.9% on MATH.”

AI News, 2025-01-20

Three things in one evening: a frontier-class result, a licence that explicitly permitted copying it, and a price roughly fortyfold below the incumbent. Any one of those makes a news cycle. Together they make something else.

Tuesday 21 January

Project Stargate: $500b datacenter (1.7% of US GDP). The largest infrastructure announcement in the corpus lands the day after, and does not displace R1 — the DeepSeek measurement stays at 135.5.

Wednesday 22 January

Bespoke-Stratos + Sky-T1: The Vicuna+Alpaca moment for reasoning. Two days after release, independent groups have distilled R1's reasoning into small models and published them. The headline's comparison is to the week in 2023 when Llama's weights leaked and the open-source ecosystem materialised in days.

Thursday 23 January

OpenAI launches Operator, its first Agent. OpenAI's biggest product launch of the month gets one day at the top of the newsletter, and the DeepSeek line barely notices — 85.9, still six times its January baseline.

Friday 24 January

TinyZero: Reproduce DeepSeek R1-Zero for $30. Four days after release, the mechanism has been reproduced at toy scale for the price of a large pizza — not R1's results, but the self-verification behaviour emerging in a small model trained from scratch, which was the part people doubted.

Monday 27 January

DeepSeek #1 on US App Store, Nvidia stock tanks −17%. Day seven. The consumer app tops the charts and the market reprices the assumption that frontier capability requires frontier capital expenditure. Community density 205.9; announcement space 376.4, and 500.0 the following day — its highest value of the entire corpus.

Figure 9 · One week, by the dayMentions of DeepSeek or R1 per 10⁴ words inside the Discord recap, by the day each issue covers, 6 January to 19 February 2025. Gaps are weekends.
05010015020001-0601-0901-1401-1701-2201-2701-3002-0402-0702-1202-1702-19R1 shipsdistilled reproductions$30 reproduction#1 on the App Storementions / 10⁴ words

Look at where the vertical markers sit relative to the line. The release moves the measurement immediately and the reproductions keep it high, but the market is the last surface to find out: Nvidia repriced on day seven, after the technical surfaces had been saturated for a full working week. Anyone reading practitioner forums knew on the Monday.

What made this different

Plenty of models in this archive match a frontier model on benchmarks. Several did it that same quarter. The R1 week is different in one specific, measurable way, and the two headlines from Wednesday and Friday are the whole of it: within four days, two independent groups had reproduced the result cheaply enough to publish, because the weights and the recipe were both in the open.

Why reproducibility is the variable

An announcement you cannot check produces one news cycle. A result anyone can reproduce produces a research programme. R1 shipped weights, a training recipe, and — through the distillation work in the same week — a path to running the capability on hardware people already owned. Each of those turns readers into participants, and participants generate more of everything: derivative models, benchmarks, tooling, arguments. That is the difference between an event and a regime change, and it is legible in the data within four days.

The decay confirms it in the opposite direction. Density falls from 215.3 on 28 January to roughly 50 by mid-February — a three-week half-life, which is ordinary. By the end of February the specific event is over.

And yet nothing went back to how it was.

The step and the spike

Figure 10 · The step and the spikeCommunity space, monthly. The company reverts below its pre-R1 level; the category it belongs to does not. Coverage ends March 2026 with the Discord recap.
020406024-0824-1125-0225-0525-0825-1126-02China blocDeepSeekmentions / 10⁴ words

These are the same measurement in the same surface over twenty months, and they separate completely.

DeepSeek — the company — spikes from 3 to 58 and then decays for a year, ending at 2 in March 2026: below where it started, before the release that made it famous.

The Chinese open-weights bloc — Qwen, DeepSeek, Kimi, GLM, MiniMax taken together — spikes to 73 and then settles into a band between 20 and 49 and stays there for fourteen months. It never returns to its pre-R1 level of 3 to 14.

Read those two lines together and the finding is this: the breakthrough permanently relocated a share of the field's attention, and gave almost none of it to the company that caused it. DeepSeek proved the category was worth watching, and then the category absorbed the gain.

It reproduces in all three surfaces, which is the minimum standard before believing anything measured this way.

Table 3 · Mentions per 10⁴ words, 2024Q4 → 2025Q1 → latest quarter
SurfaceWhatBefore R1PeakLatestOutcome
Announcement (Twitter)China bloc176572holds
Announcement (Twitter)DeepSeek alone165210reverts
Community (Discord)China bloc135044holds
Community (Discord)DeepSeek alone7417reverts
Practice (Reddit)China bloc408798holds
Practice (Reddit)DeepSeek alone177326reverts

Practice space is the most striking column. Practitioners were at 40 before R1 and are at 98 at the end of the corpus, more than a year after the week this chapter is about — the highest value the bloc reaches anywhere. Announcement space follows the same shape one step behind. On anything you can download, the people running it lead the people announcing it.

A spike tells you something happened. A step tells you something changed. They look identical for about three weeks.

What this is worth on a Monday morning

The practical content of this chapter is a test you can run on any breakthrough, in any field, without an archive:

Event or regime change?

Can other people reproduce it, and how fast? Not "is it impressive" but "is the recipe in the open." Four days is a regime change; a paper with no weights and no code is a news cycle.
Does the attention transfer to the category or stay with the author? Track the competitors, not the company. If the whole category steps up and holds, something real moved.
Where did the last surface find out? If the money moved a week after the practitioners did, that gap is the size of the edge available to anyone reading the right layer.

One more thing, and it belongs here because it is uncomfortable. On 21 November 2024, two months before this week, the newsletter's headline was DeepSeek-R1 claims to beat o1-preview AND will be open sourced. The claim was public, specific and correct. Everything in this chapter was foreseeable to anyone who read that sentence and believed it.

Almost nobody did, and the reason is not stupidity. That headline sat in an issue alongside Nvidia's quarterly revenue, a new benchmark, and four other model claims, most of which came to nothing. Correct predictions in this archive are not marked. They look exactly like the incorrect ones standing next to them, which is why the week they resolve is worth studying closely.