Chapter 3
What did the field think the job was?
In the first half of 2024, fine-tuning appears 25.6 times per
ten thousand words of the newsletter's Twitter recap, and retrieval-augmented generation runs
higher still at 30.9 — the two densest themes in announcement space that half-year. Together
that is a mention every hundred and eighty words, every day, for six months. Between them they are the
architecture: take an open model, tune it on your data, put a vector database in front of it,
ship.
Two and a half years later, in announcement space, fine-tuning is at 1.9 and RAG is at 0.2. The default architecture of 2024 essentially stopped being discussed.
The obvious reading — the field tried fine-tuning and RAG, and they did not work — is wrong for both, and wrong in two different ways. This chapter is about the world before any of that happened — what the job looked like when the answer to almost every question was take an open model and change it.
Three words carry most of that architecture, and they are worth having straight, because they are still how a model gets adapted today even where nobody says so.
Fine-tuning is continuing to train a model that someone else already trained, on data of your own, so that it behaves the way you want. The base model supplies language and world knowledge; your data supplies the job. The catch is cost: a full fine-tune updates every weight, which needs enough memory to hold the model, its gradients and the optimiser's bookkeeping — in practice several times the model's own size.
LoRA is the trick that made this affordable, and the newsletter's Discord recap gives it its full name in January 2024:
Despite the emergence of newer models, @merve3234 advocated for the effectiveness of BERT in sequence classification tasks, suggesting Low-Ranking Adaptation (LoRA) for fine-tuning to enhance parameter efficiency.
AI News, Discord recap, 2024-01-23
Instead of updating a weight matrix directly, LoRA freezes it and learns a small pair of matrices whose product is added on top. The pair has a low rank — a few million numbers standing in for a few billion — so training touches a fraction of a percent of the parameters. What you ship is not a new model but an adapter file measured in megabytes, and you can keep several and swap between them.
Quantization is the other half. Weights are normally 16-bit numbers; quantization stores them in 8, 5 or 4 bits instead, trading a little accuracy for a large reduction in memory. Combine the two — quantize the base model, train a LoRA on top — and you get QLoRA, which is why the same recap is full of people fitting training runs onto consumer cards:
Training DPO models, specifically on a 12GB graphics card, may require more VRAM than is available … Recommendations included utilizing QLoRA for fine-tuning to conserve VRAM.
AI News, Discord recap, 2024-01-23
By mid-2024 the practice layer had turned this into a menu. A single release arrives as a dozen files, each a different point on the same tradeoff:
Llama-3.1 8B Instruct GGUF models have been released, offering various quantization levels including Q2K, Q3KS, Q3KM, Q40, Q4KS, Q4KM, Q50, Q5KS, Q5KM, Q6K, and Q80. These quantized versions provide options for different trade-offs between model size and performance, allowing users to choose the most suitable version for their specific use case and hardware constraints.
AI News, Reddit recap, 2024-07-24
GGUF is the file format those weights ship in; the Q4 and
Q5 are bits per weight. None of this is glamorous and all of it is the actual
substance of what people were doing, which one commenter put better than any specification
could:
the knowledge of the whole world in a few GB of a gguf file.
AI News, Reddit recap, 2024-07-12
Mistral released the weights for an eight-expert mixture-of-experts model by posting a magnet link — no paper, no code, no inference implementation. Two days later the newsletter opened with this:
Happy Friday. 3 new models are the talk of the town today: Mistral's new 8x7B MoE model (aka “Mixtral”) — a classical attention model, done well … Mamba models, a range of models up to 3B by Tri Dao of Together … StripedHyena 7B — a descendant of the subquadratic attention replacement Hyena out of Stanford's Hazy Research lab … that is finally competitive with Llama-2, Yi, and Mistral 7B.
This is all very substantial and shows what happens when you ship model weights instead of heavily edited marketing videos.
AI News, 2023-12-08
That last line is the field's value system in one sentence, written two days after a much-criticised launch video from a much larger company. It is also a fair summary of what the next eighteen months rewarded.
The headlines that week run: The Mixtral Rush on the 9th, describing independent groups racing to write the inference code from scratch overnight so anyone could run the thing; Mixtral beats GPT3.5 and Llama2-70B on the 11th; Mixtral-Instruct beats Gemini Pro on the 15th. Seven days from an unexplained file to a model beating the previous generation's frontier, and a community whose reflex was to reimplement rather than wait.
A dense model runs every parameter for every token. A mixture-of-experts model has many parallel sub-networks and a small router that picks two of them per token, so a model with 47 billion parameters does the work of about 13 billion at inference time. You pay in memory — all the experts have to be resident — and save in compute. That trade is why MoE became the default frontier architecture, and why it arrived first in a community that had more VRAM than patience.
Read the titles from early 2024 in sequence and what you notice is that nearly all of them are about modifying models rather than using them:
LlaMA Pro — an alternative to PEFT/RAG??
AI News headlines, January–February 2024
Mixing Experts vs Merging Models
TIES-Merging
Help crowdsource function calling datasets
FSDP+QLoRA: the Answer to 70b-scale AI for desktop class GPUs
The Dissection of Smaug (72B)
Each of those is a technique for getting more out of weights you did not train. LoRA — low-rank adaptation — freezes the original model and trains a small pair of matrices alongside it, so a consumer GPU can adapt a 70-billion-parameter model without holding 70 billion gradients. Quantization stores those weights at four or five bits instead of sixteen, which is what put large models on desktop hardware at all. Model merging averages the weights of two separately fine-tuned models and, unreasonably, often produces something better than either. Synthetic data generation uses a strong model to write the training set for a weaker one.
None of these are frontier-lab techniques. They are all things you do when you cannot train a model but you can rent a GPU for an afternoon, and in 2024 that described nearly everyone who was building. The field's centre of gravity sat with people adapting other people's weights, and the newsletter's coverage sat there with them.
The most concentrated presence in the entire archive belongs to a company most engineers would now struggle to place. Mistral is tagged as a subject of 42% of all issues in late 2023 and 40% in the first half of 2024 — not merely mentioned somewhere, but one of the things the issue was about, two of every five days, for half a year. No other company outside OpenAI holds that share for that long.
By the first half of 2026 it is in 2%. It is still named in half of them; chapter 7 takes apart the difference, which turns out to be most of what losing a lead looks like.
Nothing in this chapter's period would have let you predict that, and the naive explanation — they got worse — is not what the archive says. In December 2025, the month its density reached the floor, Mistral raised $1.7 billion at an $11.7 billion valuation and shipped a coding model that practitioners reported beating or tying DeepSeek v3.2 in 71% of third-party preferences. Ceasing to be news and ceasing to be good are different events with different causes, and a chart of attention only ever shows you the first one.
The figure above is the whole period in outline, measured in announcement space: two themes that own the beginning and are near zero by the end, and one that starts at nothing and takes over. What it cannot show — because a chart of what people talked about never can — is that the two falling lines fell for opposite reasons.
Retrieval fell, on the reading this chapter argues for, because it won. That reading is an inference from the vocabulary, not a measurement of products: what the archive shows is the name going quiet while the machinery keeps being described. It is now unremarkable enough to go unnamed in most of what the newsletter covers, and a paragraph in every system design, and things in that position stop being announced. Fine-tuning fell out of announcement space while remaining one of the largest sustained practical activities in the archive: the busiest single community anywhere in this corpus, at 302,248 messages, is a fine-tuning toolchain, and it is busiest during exactly the years its coverage was collapsing.
Both fell about 90%. One is absorption and one is a coverage artifact, and no amount of looking at the falling line will tell you which is which.
A line going down is not a verdict. It is a question about where the thing went.
The answer is in the archive, in a surface the falling line does not cover. In April 2026, with fine-tuning language in announcement space down to 5.0 mentions per ten thousand words, the Reddit recap carried this:
The Chaperone-Thinking-LQ-1.0 model achieves an impressive 84% on MedQA, which is notable given its size and efficiency. This performance is achieved with a 4-bit GPTQ + QLoRA fine-tuning of the DeepSeek-R1-32B model, allowing it to run on approximately 20GB of VRAM. This makes it feasible to run on consumer-grade GPUs like the NVIDIA 3090, which is a significant advantage for accessibility and experimentation.
AI News, Reddit recap, 2026-04-21
Every element of the 2024 recipe is present — take an open model, quantize it, adapt it with a low-rank update, run it on hardware you own. What has changed is that nobody thinks this is news. It is a person reporting a medical-exam score from a bedroom GPU, and the notable thing, to the room, is the VRAM figure.
Two months later the same surface supplies the epitaph, in a single adverb:
Another commenter initially reacted to the benchmark screenshot skeptically, but updated after realizing it is Cohere's own architecture rather than merely a finetune, making the reported results more technically notable.
AI News, Reddit recap, 2026-06-11
In early 2024 a finetune of Qwen was worth a headline and a joke about the year of the Dragon. By mid-2026 merely a finetune is a reason to discount a result before reading it. The technique did not fail and did not disappear. It was demoted — from the thing you announce to the thing you are assumed to have done already, which is the same destination retrieval reached by a different road.
The editor marks the moment himself, in October 2025, and what he is describing is a technique that has begun to need defending:
the same day Jeremy Bernstein publishes Modular Manifolds … and 3 days later John Schulman writes LoRA Without Regret, an empirical endorsement of the original 2021 paper validating its performance to full finetuning similar to Biderman et al when done right.
AI News lede, 2025-10-01
Nobody writes without regret about a method that is winning. The title concedes the thing the density series had already recorded: by 2025 you had to argue for low-rank adaptation, where in 2024 you simply used it.