Written ForwardsChapter 10

Chapter 10

How ideas die

Four ideas ended and one only looked like it. How do you tell?

On 8 December 2023 the newsletter ran three new models side by side, under the headline Mamba v Mistral v Hyena, and treated them as live competitors on equal footing:

Mistral's new 8x7B MoE model (aka “Mixtral”) — a classical attention model, done well … Mamba models, a range of models up to 3B by Tri Dao of Together … StripedHyena 7B — a descendant of the subquadratic attention replacement Hyena out of Stanford's Hazy Research lab … that is finally competitive with Llama-2, Yi, and Mistral 7B.

AI News, 2023-12-08

One of those three is now inside nearly every model you use. Of the other two, one is a footnote and one you have probably never heard of.

The interesting part is not that two of them lost. It is that the archive contains four different ways to lose, they look almost identical in a chart, and only one of them means the idea was wrong. Getting the distinction right is the difference between correctly dropping a dead technique and abandoning something that was merely early.

The bet against the transformer

Attention — the mechanism at the heart of every transformer — compares every token in the input with every other token. That is quadratic: double the context and you quadruple the work. In late 2023 this looked like the wall the whole field would eventually hit, and there was a serious, credentialed alternative.

State-space models, as the field met them

A state-space model processes a sequence the way a control system does: it keeps a fixed-size internal state, updates it one token at a time, and never looks back at the raw history. Cost is linear in sequence length instead of quadratic, and inference needs no growing key-value cache. Mamba, RWKV and StripedHyena were three takes on the idea. The trade is that a fixed state must forget things, and what it forgets turns out to matter for exactly the recall-heavy tasks people use language models for.

In the first half of 2024, state-space language runs at 12.13 mentions per ten thousand words of announcement space. Mixture-of-experts — the architecture that did win — runs at 12.83 in the same window. They were, by this measure, equally live questions.

In the last half-year of the corpus, state-space language appears zero times in 46,815 words of announcement space. It falls in practice space too (1.62 to 0.39) and in community space (2.38 to 0.43). It happened in every surface at once, which rules out the dullest explanation — that the newsletter simply changed what it was sampling. Something real ended.

Figure 5 · Four ideas, announcement spaceMentions per 10⁴ words inside the Twitter recap. Three of these lines end at or near zero and one is climbing — but the shape of a line says nothing about which fate produced it.
048122024H12024H22025H12025H22026H12026H2mentions / 10⁴ wordsworld models1-bit / ternarymodel mergingstate-space

Fate one: absorbed below the vocabulary

Here is the thing that makes the naive reading wrong. Search the archive for the day state-space models were declared dead and there is no such day. What there is instead is a headline from 13 June 2024:

Hybrid SSM/Transformers > Pure SSMs/Pure Transformers

AI News, 2024-06-13

And then, eighteen months later:

NVIDIA Nemotron 3: hybrid Mamba-Transformer completely open source models from 30B to 500B

AI News, 2025-12-15

A 30-billion-parameter open-weights model with a one-million-token context window, built from Mamba layers interleaved with attention layers, shipped with weights, training recipes and datasets. The pure-SSM bet lost. What this establishes is narrower than “state-space models won”: it is that the mechanism survived, in a hybrid, in at least one shipped family. One model line is not evidence of broad absorption, and the corpus cannot tell you how widely those weights were used.

What did not happen is the part worth noticing: no vocabulary rose to replace the one that fell. Hybrid-architecture language never picks up the slack — it sits at one to three mentions per ten thousand words in every surface across the whole corpus, on counts small enough (never more than 29 in a half-year) that the series is mostly noise. The architecture stopped being a topic and became an implementation detail, which is what winning quietly looks like.

An idea that gets absorbed leaves the same hole in the record as one that failed. The difference is whether a product exists that nobody argues about.

Fate two: the actual death

Model merging is the control case. The technique: take two models fine-tuned separately from the same base, average their weights, and — unreasonably — often get something better than either. It had tooling (MergeKit), named methods (TIES-merging, SLERP, task arithmetic) and a genre of results (frankenmerges — models stapled together out of duplicated layers). It peaks in announcement space in the second half of 2024 at 2.59.

Then: 9 mentions, 2 mentions, 0, 0. Practice space ends at zero in the same window. Community space falls from 199 mentions in a half-year to 43.

No successor headline. No hybrid. No product that quietly contains it. This is what a real death looks like in this data. The reason is legible from the archive around it: merging was a technique for a world in which everyone had a pile of their own fine-tunes to combine — and over the same window, fine-tuning language in announcement space fell from 25.6 to 1.9. When that world ended, the technique had nothing left to operate on.

Fate three: refuted, then revived from below

On 1 March 2024 the headline was The Era of 1-bit LLMs. The paper behind it — BitNet b1.58 — proposed training models whose weights are restricted to three values, −1, 0 and 1, which removes multiplication from the forward pass almost entirely. Announcement-space density hits 6.96. It is, briefly, the most exciting idea in the corpus.

The reason it stalled is in the archive too, and it is a result rather than a mood. On 12 November 2024 the newsletter led with a paper that measured the trade directly:

A group of grad students under Chris Ré has now modified Chinchilla scaling laws for quantization over 465+ pretraining runs and found that the benefits level off at FP6 … the longer you train, the more data seen during pretraining, the more sensitive the model becomes to quantization at inference-time … this loss degradation is roughly a power law in the token/parameter ratio seen during pretraining.

AI News, 2024-11-12

In other words: the harder you train a model, the less of it you can throw away afterwards — and everyone was training harder every month. The headline the editor put on that same issue was blunter than the finding.

BitNet was a lie?

AI News, 2024-11-12 — the headline over the passage above

By the first half of 2025, the density is 0.08 — one mention in 131,271 words of announcement space. That is as dead as anything in this archive gets.

And then, in July 2026, it comes back — in the wrong surface. The newsletter's Reddit recap fills up with a ternary variant of Qwen3.6 27B, reported as compressing roughly 54GB of weights to under 4GB and running locally in a browser through custom WebGPU kernels. Practice-space density reaches 4.71, its highest value in the corpus, against 1.71 in announcement space. For the first time in the life of this idea, the people running it are ahead of the people announcing it.

The reception is also exactly what practice space is for:

Commenters pushed back on the wording “near fp16 precision” because ternary weights are {−1, 0, 1} … A commenter distinguished 1-bit models trained from scratch from extreme post-training quantization, arguing the former should retain much more capability … Several asked for rigorous benchmarks such as SciCode or SWE-rebench.

AI News, Reddit recap, 2026-07-15

Nobody there is excited about the era of anything. They are asking which 4GB model beats which other 4GB model on a laptop they own. That is the signature of an idea that was not wrong, only early: it was waiting on kernels and hardware rather than on insight, and it reappeared the moment someone shipped the kernels.

Figure 6 · An idea returning from belowThe same pattern in the Twitter and Reddit recaps. Announcement space peaks in 2024 and effectively stops; practice space reaches its corpus high two years later, on hardware that did not exist for it in 2024.
02462024H12024H22025H12025H22026H12026H2mentions / 10⁴ words1-bit (practice)1-bit (announcement)

Fate four: still open

On 10 March 2026 the headline was Yann LeCun's AMI Labs launches with a $1.03B seed to build world models. World-model language in announcement space starts at 0.31 in early 2024, sits under one mention per ten thousand words for a year and a half, and then climbs to 4.59 by early 2026 and 5.34 in the last half-year. That is a seventeenfold rise, which is precisely the sort of ratio this book has spent two interludes telling you not to trust: the denominator is four mentions in 47,175 words. The levels at the end are the trustworthy part, and they are large.

Now compare the surfaces. The practice-space baseline is a single mention, so a fold change is meaningless; compare levels instead. In the first half of 2026 world models run at 4.59 in announcement space, 1.22 in community space, and 0.29 in practice space — a descending staircase from the people announcing things to the people running them, which in this book is the signature of a narrative rather than an adoption.

Except this time I do not think the test applies, and it is worth being precise about why.

Where the practice surface is blind

Practice space is people running models on hardware they own. That makes it an excellent check on anything downloadable and no check at all on anything requiring a data centre. You could not run a world model in 2026 if you wanted to. Reasoning models had the same problem when they arrived in late 2024 — nearly invisible in practice space for months — and they were entirely real. When the practice surface is quiet because a thing is unrunnable rather than uninteresting, its quiet carries no information.

So the honest verdict on world models is: undecided, and the instrument that settled the other three cases cannot settle this one. That is a less satisfying answer than a prediction and a more useful one, because it tells you what evidence would change your mind — the first world model somebody can run on a consumer GPU, and what practice space says about it the following week.

Four fates, four signatures

Table 1 · Mentions per 10⁴ words, 2024H1 → 2026H2, in each surface
IdeaAnnouncementPracticeFate
state-space models12.13 → 0.001.62 → 0.39absorbed (below the vocabulary)
model merging1.08 → 0.000.97 → 0.00gone
1-bit / ternary6.96 → 1.711.94 → 4.71deferred, revived in practice
mixture-of-experts12.83 → 7.694.85 → 11.08absorbed (into infrastructure)
world models0.31 → 5.340.32 → 0.10undecided

None of these four is visible from the falling line alone. A chart of attention tells you where the conversation went; it never tells you why, and the why is the entire decision. What it does give you, if you look at more than one surface and read a few of the days around the break, is enough to sort a fall into the right bucket — which is all the decision usually needs. The rules for doing that sorting are worth stating once, at the end of the chapter, when there is a fifth fate to put beside them.

A fifth case, and the one that breaks the scheme

Four fates, four signatures, and a scheme that looks complete. It is not, and the idea that breaks it is the one this book has already used twice — because the largest fall in the whole archive belongs to none of the four.

Retrieval-augmented generation is the most complete disappearance in this archive. Inside announcement space the name runs at 22.6 mentions per ten thousand words in early 2024 — the densest technical idea in the corpus at that point — and in the final half-year it runs at 0.21. That is a hundred-and-sevenfold fall. Nothing else measured in this book falls that far.

One note on which pattern that is, because this chapter is about to compare a name against a mechanism and the comparison is worthless if the two are counted on different footings. The 30.9 quoted back in chapter 3 is a broader pattern that also catches retrieval-augmented written out in full; it falls to 0.2, the same shape and a little larger. Everything in this chapter uses the narrow one — the string people actually said when they meant the architecture — so that the name and its machinery are measured the same way.

What retrieval augmentation actually was

The idea is one sentence long. A model knows only what was in its training data, and its context window — the text you can hand it at question time — is finite and expensive. So instead of retraining the model on your documents, you fetch the handful of passages that look relevant to the question and paste them into the prompt. The model does not learn your data; it reads it, once, per question.

Making that work is a pipeline, and every stage is a choice. You chunk documents into passages, because a whole PDF will not fit. You embed each chunk — run it through a model that turns text into a vector, so that passages about the same subject land near each other. You put those vectors in an index so you can find the nearest ones quickly. At question time you embed the question, retrieve its neighbours, and prepend them.

The failure modes are all in the choices, and the archive is full of them from the start. February 2024, a complaint about the retrieval built into ChatGPT's custom GPTs:

Nick Dobos (of Grimoire fame) also blasted the entire knowledge files capability — it seems the RAG system naively includes 40k characters' worth of context from docs every time, reducing available context and adherence to system prompts.

AI News lede, 2024-02-01

That is the canonical failure in miniature. Retrieve too much and you crowd out the instructions, spend money on tokens nobody reads, and make the model worse at following its own system prompt. Retrieve too little, or chunk badly, and the answer is not in the context at all. The engineering is entirely in the middle.

Two years later the same problem is still being worked, at a scale that makes the tradeoffs explicit:

LEANN: “stop storing embeddings”. A notable systems claim: index 60M text chunks using 6GB (vs “200GB”) by storing a compact graph and recomputing embeddings selectively at query time … Engineers should sanity-check latency/throughput tradeoffs and recall under recomputation.

AI News, Twitter recap, 2026-01-07

Sixty million chunks, a storage budget, and recall as the thing you might lose — recall being the share of genuinely relevant passages your index actually returns. Those are the same four words as 2024: chunk, embed, index, retrieve. Only the scale and the storage layout changed.

The name and the machinery

A fall that steep has two possible explanations and the line cannot tell them apart. Either the technique failed, or it succeeded so completely that people stopped naming it. The way to find out is to stop counting the name and start counting the mechanism — embeddings, chunking, reranking, vector indexes, semantic search, BM25 — the words you have to use if you are actually building retrieval, whether or not you call it RAG.

Half-year“RAG”, the name its machinerymachinery ÷ name
2024H122.614.20.6×
2024H221.112.40.6×
2025H15.64.50.8×
2025H25.612.82.3×
2026H12.613.45.2×
2026H20.27.535×

In 2024 the name is mentioned rather more than the mechanism: if you were talking about embeddings you were probably saying RAG in the same breath, and often only that. The two then come apart, and the last column does it monotonically — every half-year the machinery gains on the name, and it never once gives ground. The name falls a hundred and sevenfold. The machinery falls 1.9-fold, and in the second half of 2025 it is higher than it was in the second half of 2024 — more talk about chunking and reranking, at the moment the word RAG had all but vanished.

The name fell a hundredfold. The thing it named fell by half, and spent a year going back up.

This is what absorption looks like from inside a text corpus, and it is worth being precise about what it does and does not establish. It does not show that retrieval was deployed more widely; nothing in this archive can show that. It shows that the vocabulary of building retrieval outlived the label by a factor that grew every half-year for two years — which is the signature of a technique that stopped being a topic and became a component.

Put the chapter's two quotations side by side and the reclassification is visible without any counting at all. In 2024, Nick Dobos is blasting the RAG system by name, as a product surface that behaves badly. In 2026, the same subject arrives as LEANN — sixty million chunks, a storage budget, a recall tradeoff — and the name survives only as a modifier, local RAG, attached to the thing that actually matters, which is the graph. Two years apart: a feature with a name, then a storage problem with a benchmark. The word did almost all of the dying.

"RAG is dead" was a real position, argued in public by people who build things, and the data above is exactly what you would show to support it.

It is also wrong, and the second column is what shows it. Reading a few weeks of coverage around the fall would have shown it too — but reading does not scale and it is not falsifiable, and the second column is both. Stated generally, so it works on something other than retrieval:

The name / machinery test

Measure two things instead of one. The name — what the idea is called. And the machinery — the vocabulary of the mechanism the idea needs in order to work at all, which for RAG means retrieval, chunking, reranking, vector indexes, embeddings, BM25, hybrid search.

If the name falls and the machinery holds, the idea won so completely that naming it became unnecessary. If the name falls and the machinery falls with it, the idea died. The machinery is the tell, because an idea that shipped keeps generating machinery talk — things inside working systems still break, get tuned, and get argued about.

Figure 19 · The name died; the machinery did notAnnouncement space. RAG's own name falls 107-fold. The vocabulary of the mechanism it needs — retrieval, chunking, reranking, vector indexes, embeddings — falls 1.9-fold, and memory language rises.
010202024H12024H22025H12025H22026H12026H2mentions / 10⁴ wordsmemoryretrieval machineryRAG (the name)

The two lower lines are the table above, drawn. The line above both of them is the one the table does not contain, and it is the one that settles the case. Memory — long-term memory, memory layers, what a system keeps and retrieves across turns — goes from 12.91 to a peak of 20.74. The job RAG existed to do is discussed more at the end of the corpus than at the beginning. It simply is not called RAG, because it stopped being an architecture you choose and became something the software around the model does on every turn.

You can watch that reclassification happen in the archive's own prose. In 2024, retrieval is the subject of the sentence. By mid-2025 it has become one item in a list of ingredients — here is a practitioner definition quoted in the newsletter in June 2025, of the thing that replaced it: “filling the context window with just the right information for the next step … task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history, compacting.” RAG appears in that sentence as a comma-separated component of somebody else's job.

The control

A test that returns “absorbed” for everything is worthless, so run it on something that actually died.

Model merging — averaging the weights of two separately fine-tuned models to get something better than either — had tooling, named methods and a genre of results in 2024. Its name goes 1.16, 2.40, 0.69, 0.14, 0.00, 0.00. Its machinery — weight averaging, weight interpolation, task vectors, task arithmetic — never exceeds 0.37 anywhere in the corpus, in any surface, and is 0.00 at the end. There is no residue. Nobody argues about the tuning of a thing nobody ships.

That is what death looks like, and it does not look like RAG.

Three more, without commentary

Table 6 · Announcement space, 2024H1 → 2026H2, mentions per 10⁴ words
IdeaThe nameThe machineryVerdict
RAG22.57 → 0.2114.22 → 7.48absorbed
fine-tuning26.35 → 1.5016.07 → 9.83absorbed
prompt engineering1.78 → 0.211.55 → 2.56absorbed, machinery grew
MMLU / HumanEval / GSM8K4.33 → 0.210.77 → 9.40replaced
model merging1.16 → 0.000.00 → 0.00dead

Read the middle two columns as one ratio. Fine-tuning is the case most often cited as a technology that failed: the name falls eighteenfold. Its machinery — LoRA, QLoRA, PEFT, adapters, SFT, instruction tuning, post-training, synthetic data — falls 1.6-fold and sits at 9.83 at the end of the corpus. Fine-tuning did not stop. It stopped being a topic and became a step.

Prompt engineering is the strongest case in the table: the name falls eightfold while its machinery — system prompts, few-shot examples, instruction files, compaction, context management — rises 1.7-fold. An idea whose vocabulary of practice grows while its name disappears is not in decline by any reading.

The fifth fate

The fourth row is doing something the other rows are not, and it needs its own name.

MMLU, HumanEval and GSM8K — the benchmarks that defined 2024 — fall twenty-onefold, from 4.33 to 0.21. But their machinery is not diffuse mechanism vocabulary. It is a specific list of successors: SWE-bench, ARC-AGI, GPQA, FrontierMath, Terminal-Bench, SciCode, AIME, LiveBench. Those rise from 0.77 to 9.40, a twelvefold gain, and they occupy the exact role the old ones did.

That is replacement, and it is worth separating from absorption because the two imply opposite actions. When something is absorbed, the mechanism is still there and your system is already using it; leave it alone. When something is replaced, there is a named successor doing the same job and you should migrate.

Five fates, and how to tell them apart from the outside
FateThe nameThe machineryWhat to do
Absorbedfalls hardholds or risesnothing; you are already using it
Replacedfalls harda named successor risesmigrate
Deadfalls to zerofalls to zerodrop it, and check what depended on it
Deferredfalls in announcement, returns in practicereappears with new hardware or toolingwatch practice space
Artifactfalls in the aggregate onlyholds in every surfacefix the instrument

When the test fails

It is a cheap test and cheap tests have failure modes. Three of them matter.

Generic machinery vocabulary. My memory line is the weakest number in this chapter: the word means at least three things in this corpus — what a system retains across turns, what a model has memorised, and how many gigabytes a GPU has. The rise from 12.91 to 20.74 is real but it is not purely about retrieval, and I would not build a decision on that line alone. The retrieval-machinery line, which uses specific terms, is the one carrying the argument.

Shared machinery. Two ideas can need the same mechanism, in which case the mechanism's persistence tells you one of them survived but not which. Embeddings serve retrieval, classification, clustering and search alike.

The instrument. All of this has to be measured inside a fixed source. The mix of sources in this archive inverts over three years, and a count taken across the whole document partly measures that inversion rather than the field — so run the name/machinery test on unsegmented text and you will get a confident answer about your own sampling.

Running it forwards

The test also works in the other direction, which is where it earns its keep, because falling lines are a retrospective problem and rising lines are a decision you have to make now.

Take the biggest rising line in the archive. Agentic language in announcement space rises eightfold between 2024 and 2026. Its machinery — orchestration, sub-agents, tool-use, sandboxing, scaffolding — rises nineteenfold, faster than the name itself. Whatever else is true about the agent narrative, real engineering vocabulary is accumulating underneath it faster than the label is, which is not what a purely marketed term looks like. Compare model merging in its best year: the name rose and the machinery never arrived at all.

So the rule is symmetric, and it is the whole chapter in two lines:

A name that rises faster than its machinery is a term being marketed. A name that falls while its machinery holds is a technology that won.

Neither is visible in the line everybody quotes, and both take about twenty minutes to check.