Chapter 7
How does a technological lead change hands?
Here are two measurements of the same company in the same year. Mistral is named somewhere in 48% of the issues published in 2026 — every other day. Mistral's density inside the newsletter's announcement recap in 2026 is 1.3 mentions per ten thousand words, down from 15.2 two years earlier.
Both are correct. Together they describe what losing a technological lead actually looks like from the inside, which is not disappearance. It is becoming background.
In early 2024 the open-weights frontier had two names on it, and between them they were the loudest thing in the archive. Meta and Mistral together run at 42.3 mentions per ten thousand words of announcement space and 65.5 in practice space — higher than any single lab reaches at any point in the corpus. By the end they are at 1.3 and 2.4.
Separate them and the two falls have completely different shapes.
Meta breaks. Its line does not decline from 2024; it rises to a peak of 34.5 in the second half of 2024 — Llama 3.1 and the 405B, the most-discussed open-weights moment in the corpus to that point — and then falls off a cliff: 15.4, 4.3, 1.4, 1.1. Thirty-one-fold from peak in two years.
Mistral erodes. 15.2, 10.3, 8.7, 7.9, 2.3, 1.3. No peak, no cliff, no event. Every half-year is a little lower than the one before, for two and a half years, until there is nothing left.
On 19 April 2024 the headline is Meta Llama 3 (8B, 70B), and the next day Llama-3-70b is GPT-4-level Open Model. Here is the lede from that second issue, which is the high-water mark of American open weights and reads like it:
With a sample size of 1600 votes, the early results from Lmsys were even better than reported benchmarks suggested, which is rare these days … This is the first open model to beat Opus, which itself was the first model to briefly beat GPT4 Turbo. Of course this may drift over time, but things bode very well for Llama-3-400b when it drops.
Already Groq is serving the 70b model at 500–800 tok/s, which makes Llama 3 the hands down fastest GPT-4-level token source period … Llama 2 and 3 (and Mistral, to a less open extent) have pretty conclusively consigned Chinchilla laws to the dustbin of history.
AI News, 2024-04-19
Confident, specific, and correct about everything it could check. Note the forward-looking sentence in the middle — things bode very well for Llama-3-400b when it drops — which is exactly the kind of claim a retrospective would quietly leave out.
Forty-eight days later:
Qwen 2 beats Llama 3 (and we don't know how)
AI News, 2024-06-06
Read the parenthesis again. It is doing more work than the rest of the sentence. The field could see the thing happening and could not account for it — in June 2024, eighteen months before anyone would describe the handover as complete. It is the earliest actionable signal in this archive, and it is a joke in brackets.
Llama 4's Controversial Weekend Release. Two mid-size mixture-of-experts models and a promised two-trillion-parameter “behemoth”, with genuinely new engineering — early fusion with MetaCLIP, interleaved chunked attention without RoPE, native FP8 training, up to 40 trillion tokens. Released on a Saturday, and received badly. Change-point detection on the monthly Llama series puts structural breaks either side of it, in October 2024 and August 2025. It is the last time Meta's line moves at all.
And then, on 29 December 2025, this:
Meta Superintelligence Labs acquires Manus AI for over $2B, at $100M ARR, 9 months after launch
AI News, 2025-12-29
Meta's announcement-space density in that half-year is 1.4. The company was spending billions and had almost no share of the conversation. Capital and attention had come completely apart, which is worth remembering the next time either one is offered as evidence of the other.
This is the finding the chapter exists for, and it is easy to miss because everyone remembers a single name.
DeepSeek peaks at 34.5 in the first half of 2025, the half-year R1 shipped in, and then falls to 10.9, 8.6, 10.0. It never leads again. Qwen never spikes at all: 4.3, 1.8, 10.7, 11.5, 6.9, 6.4 in announcement space, and in practice space it climbs steadily from 1.3 to 22.8, the most consistent single line in this book. Kimi is at zero for two years and then 19.8, 15.7, 40.8 — the largest single-lab value anywhere in the corpus, in the final half-year. GLM arrives at 9.5 in late 2025 and holds. MiniMax comes up behind it.
No individual Chinese lab holds the top position for more than two consecutive half-years. The lead did not pass from Meta to DeepSeek. It passed from two named companies to a rotating cast of five, and the rotation is the point: whichever one happened to be ahead, the category kept the gain.
The obvious strategic response in early 2025 was “switch to DeepSeek.” That would have been a bet on the single line in this chart that reverted — from 34.5 down to 10.0 while the category around it held. The durable read was never a company. It was that a category had become viable and the names inside it would keep changing. If you are choosing a dependency, the question that survives contact with this data is which ecosystem your tooling, quantizations and fine-tunes will follow, not which lab posted the best number this quarter.
Return to Mistral, because the way the archive appears to contradict itself about it is worth more than the arc.
| How you ask | 2023 | 2024 | 2025 | 2026 |
|---|---|---|---|---|
| Named anywhere in the issue body | 100% | 99% | 87% | 48% |
| Tagged as a subject in front matter | 54% | 33% | 15% | 2% |
| Density in announcement space | — | 12.8 | 8.3 | 1.8 |
These are three different questions wearing the same clothes. Is it still around (named in the body). Is it still the story (tagged as a subject). How much of the conversation is it (density in a fixed section). Ask the first and Mistral is fine; ask the third and Mistral has essentially vanished. Both conclusions have been published, by people looking at the same archive.
And the underlying reality is stranger than either. In the same month its density hit the floor, Mistral raised $1.7 billion at an $11.7 billion valuation and shipped Mistral Large 3 plus three sizes of Ministral, all open weights under Apache 2.0. A week later practitioners were reporting that its Devstral 2 Small "beats or ties DeepSeek v3.2 in 71% of third-party preferences while being smaller, faster and cheaper." The newsletter's own lede that day was two words long:
Mistral is back!
AI News lede, 2025-12-02
The density series never registers it. Not a bump.
Being good and being the story became independent variables, and only one of them is visible in any measurement of attention.
It shows the open-weights frontier relocating, in every surface, over about eighteen months. It shows the first legible warning arriving in June 2024 and being treated as a curiosity. It shows the replacement being a bloc rather than a company. It shows a firm continuing to ship competitive models, raise enormous sums, and win third-party comparisons while its share of the conversation went to nearly nothing.
What it does not show is why — and this book will not pretend otherwise. A record of what a field discussed cannot tell you whether Meta's problem was organisational, whether Mistral's was distribution, or whether the whole thing was decided by the cost of compute in two countries. Those are questions for evidence this corpus does not contain.
What it does tell you is the shape, and the shape has a use. A lead changes hands slowly, visibly, and with the first warning arriving about eighteen months before anyone acts on it — phrased, on the day, as a joke.