Written ForwardsChapter 13

Chapter 13

Weights you can keep

What changed when a model you could download caught up?

Every other vocabulary in this book rises or falls over years. Open weights does neither for two and a half years — and then triples in the final six months, in both surfaces at once.

Open-weights language24H1 24H225H125H226H1 26H2
Announcement space11.97.49.0 11.710.730.3
Practice space5.56.85.4 5.17.733.9

By this book's own test that is the least ambiguous reading available: both surfaces move, they move together, and they move by about the same multiple. Whatever happened did not happen to the coverage.

What “open weights” means, and what it does not

The term is precise and narrower than open source. Open weights means the trained parameters are published: you can download the file, run it on your own hardware, fine-tune it, quantize it and serve it without asking anyone. It generally does not mean the training data is published, or the training code, or that the licence permits every commercial use. You are given the artefact, not the recipe.

For a working engineer that distinction decides which risks you own. A published weight file cannot be deprecated out from under you, cannot change its behaviour overnight, cannot raise its price and cannot read your inputs. It also cannot be patched by anyone else, comes with no uptime, and is entirely your problem to serve.

The long flat middle

For most of the corpus this conversation is oddly steady, and the steadiness hides the actual event, which was happening under a different name. The models practitioners were downloading increasingly came from one place:

Chinese-lab language24H1 24H225H125H226H1 26H2
Announcement space10.213.348.5 51.146.772.0
Practice space6.531.549.5 37.057.698.4

Qwen, DeepSeek, Kimi, GLM and MiniMax end the corpus at 98.4 mentions per ten thousand words of practice space — the densest any subject gets on that surface in the archive's later years, and consistently ahead of its own announcement figure. This is the clearest practice-led gradient in the book: practitioners were running these models in volume while the announcement layer was still catching up.

For most of this period “open weights” and “Chinese lab” were the same sentence, and only one of them was being counted.

What happened at the end

Then the final half-year, when the term stops being a niche and becomes the frame everyone argues inside. Two ledes give the texture:

At Computex in Taiwan, Jensen also brought the heat with Nemotron 3 Ultra, their 550B-A55B, remarkably efficient/fast open weights LLM that is the new US SoTA.

AI News lede, 2026-06-01

On any given Sunday, the announcement that the 2.4T param Qwen 3.8 Max will be open weight would've earned title story status, but they had the misfortune to do this 4 days after Kimi K3 2.8T was announced.

AI News lede, 2026-07-20

Read together, the step change explains itself. In the first, an American hardware company ships an open-weights model and the newsletter calls it the US state of the art — the category has stopped being a Chinese speciality and become a competitive position Western vendors want. In the second, a 2.4-trillion-parameter open release is not the headline, because a larger one arrived four days earlier. Open weights had become crowded enough that scale alone no longer bought attention.

Who needs the file rather than the API

The abstract case for open weights is autonomy. The concrete case, in practice space, is narrower and more convincing — it is people for whom sending the input somewhere else is not an option:

A commenter working in drug discovery noted they “can't use mainstream/closed LLMs,” implying constraints around proprietary molecular and IP data, confidentiality, compliance and auditability when sending prompts to hosted models … domains like pharma may prefer local or open-weight models to avoid data exfiltration and policy-filter limitations.

AI News, Reddit recap, 2026-06-25

That is the whole argument in one sentence, and it has nothing to do with cost or ideology. For some data the question is not which model is best, it is which models are legally reachable at all.

What it costs to exercise that option is the other half, and practice space is unsentimental about it:

Commenters focused on the deployment feasibility of unsloth/GLM-5.2-GGUF, with one noting the apparent 800GB footprint and asking how aggressively it would need to be quantized to run locally. Another technical concern was KV-cache scaling for very long context: “imagine the KV Cache size to reach 1M CTX”.

AI News, Reddit recap, 2026-06-17

Eight hundred gigabytes of weights before you have served a single token, and a memory cost for context that scales on top of that. The freedom is real and so is the bill; what changed by 2026 is that the bill was worth arguing about in public.

Why the price series moves with it

The other vocabulary rising through the same window is money — price cuts, cost per token, throughput, tokens per second:

Cost and throughput language24H1 24H225H125H226H1 26H2
Announcement space1.91.82.2 6.86.59.4
Practice space3.63.02.9 4.26.210.4

A fivefold rise in one, a near-threefold rise in the other, tracking each other closely enough to be one phenomenon. This is what the two previous sections cash out as. Once a downloadable model is within reach of the frontier, the frontier has to compete on price, and the thing everybody discusses stops being what a model can do and becomes what it costs to make it do it. Test-time compute made quality a dial the caller pays for; open weights put the floor under that price at approximately zero.

It is the least glamorous vocabulary in this book and the one whose rise is hardest to argue with. It moves in both surfaces, at the same time, in the same direction, for a reason anyone running a service can state in a sentence.

Figure 27 · Flat, then a stepTwo and a half years of a steady conversation, then a tripling in the final half-year in both surfaces at once. The money vocabulary rises through the same window.
01020302024H12024H22025H12025H22026H12026H2mentions / 10⁴ wordsopen weights (practice)open weights (announcement)cost and throughput

What that flat middle conceals is worth stating plainly, because it is the reason this chapter is short and its finding is not. For two and a half years the words open weights did not move, while the thing they describe moved continents — from Llama and Mistral to a rotating bloc of five Chinese labs, at four to ten times the density, on the surface where people actually download things. The vocabulary only caught up in the last six months, once a Western hardware company shipped an open model and called it the US state of the art. The category had been the story for two years before the category had a name anyone was using.

A term that stays flat while the world underneath it changes is not a control. It is a term nobody needed yet.