Chapter 6
When did the field stop talking about models and start talking about the software around them?
Train a word-embedding model on the archive's 2024 text and ask it what
harness is closest to. The answer comes back:
lm-evaluation-harness, eval, lm-eval,
helm. In 2024 a harness was a test runner — the thing that fed a benchmark
suite to a model and collected the scores.
Train the same model on the archive's 2026 text and ask again:
orchestration, harnesses, ux,
abstraction.
Same word. Different thing. The cosine distance between the two neighbourhoods is 0.439, the third-largest drift of any term in the corpus.
Now count it. In the whole of 2024, the word harness appears in announcement
space exactly once — one mention in the first half of the year, and none at
all in the second. In the same two years it appears 4,195 times in community
space, and 84% of those fall within ninety characters of the string
lm-eval. They are not discussions of a concept. They are a project's bug tracker:
Don't use
AI News, Discord recap, 2024-01-23get_task_dict()in task registration / initialization by haileyschoelkopf · Pull Request 1331 · EleutherAI/lm-evaluation-harness
By 2026 the word runs at 22.7 and then 31.83 per ten thousand words of announcement space, and it means this:
Harnesses may be the real differentiator: current agent harnesses underutilize frontier models; the key low-hanging fruit is turning setup into continuous skill-building — when the agent makes mistakes, it should patch itself with new skills, protections and reminders, effectively a lightweight continual-learning loop.
AI News, Twitter recap, 2026-01-02
A dashboard counting the string harness since 2024 would show a rise from one
mention to several hundred, and would be measuring two unrelated things in two different rooms.
Nothing about the series would look wrong. Worse, a whole-corpus count would show the word was
already common in 2024 — because community space was 96% of the archive's words that
year, and a single repository's pull requests were enough to make a term look established.
The word was not rare in 2024 and then popular in 2026. It was a filename in a chat room, and then a design problem on a main stage.
The drift is a warning and worth taking. But it is not the subject of this chapter. The subject is what the new meaning is for, because the word changed at exactly the moment the field's centre of gravity moved off the model and onto the software around it.
Everything that is not the model. The loop that decides to call it again; the tools it is allowed to invoke and what happens when one fails; the sandbox the whole thing runs in; what goes into the context window, when, and what gets thrown out to make room; how many attempts before giving up; when to stop. In 2024 you wrote a prompt and got a completion. In 2026 you write a harness, and the model is one component inside it — the one you did not write, cannot debug, and swap out every few months.
You can tell that list is real rather than tidy because of the arguments it produces. Here is the context-window line item, being fought over in August 2025:
Agentic coding posture: multiple reports argue for non-reasoning coders by default — reasoning can exhaust context in agent loops. Early Cline tests found v3.1 “makes assumptions” in planning; tracking diff edit failure rate as more data comes in.
AI News, Twitter recap, 2025-08-22
Every clause is a harness decision. Whether to use a reasoning model at all is a budget question, because reasoning tokens consume the same context the loop needs to remember what it has already tried. Diff edit failure rate is a harness metric — how often the model's proposed change fails to apply — and it has nothing to do with how clever the model is. This is what it looks like when the interesting engineering has moved outside the weights.
harness is not a special case. Train a word-embedding model on the archive's
2024 text and another on its 2026 text, align the two spaces, and the distance a word has moved
is measurable. Across the corpus the median term moves 0.313. The terms that move furthest are
the ones this book is about.
| Term | Drift | Nearest neighbours, 2024 | Nearest neighbours, 2026 |
|---|---|---|---|
skills | 0.590 | goals, abilities, experiences | middleware, reusable, ide, filesystem |
distillation | 0.448 | unet, dare, neuron, imagenet | attacks, industrial-scale, copyrighted, laws |
harness | 0.439 | lm-evaluation-harness, eval, lm-eval, helm | orchestration, harnesses, ux, abstraction |
agentic | 0.371 | empowers, augmenting, production-ready, devika, low-code | long-horizon, tool-use, multi-step, computer-use, swe |
prompt | 0.363 | prompts, engineering, crafting, promptfoo | injection, adherence, caching, drafting |
reasoning | 0.285 | multi-hop, chain-of-thought, abilities | multilingual, spatial, instruction-following |
safety | 0.272 | disclosure, copyright, legal, risk, regulatory | political, mass, surveillance, misuse |
context | 0.170 | window, length, contexts, yarn | window, length, cache, kv, batch |
Read the two right-hand columns as a pair and the pattern is not that these words got more
advanced. It is that they changed subject. Skills moves furthest of anything
measured, 0.590: in 2024 its neighbours are goals, abilities,
experiences — the vocabulary of what a model can do. By 2026 they are
middleware, reusable, ide, filesystem. The word stopped
describing a capability and started describing a file on disk that an agent loads.
Two of the low-drift rows are the control. Context barely moves at 0.170 —
window, length, cache — and reasoning at 0.285, which is the more interesting one:
the amount of discussion about reasoning changed enormously, as the chapter on it
shows, while the word's meaning barely shifted at all. Volume and meaning are independent, and
a series that tracks one tells you nothing about the other.
Re-derive what your metrics mean, on a schedule. A dashboard is a set of strings frozen on the day you wrote it, and the field does not consult it before changing vocabulary.
The whole layer moves together. Coding agents rise twenty-onefold from 2.55 to a peak of 54.60. Orchestration language — multi-agent, sub-agent, scaffolding, agent loops — rises twenty-threefold. Sandboxing, which is what you need once a model is running code you did not read, rises from 1.24 to 10.68.
And one line goes the other way.
Prompt engineering falls from 2.86 to 0.43 in announcement space, and from 2.26 to 0.39 among practitioners. In early 2024 it was a named skill with conference talks and job titles. By the end of the corpus it is a rounding error in both surfaces, which by the chapter-2 test means it is genuinely gone rather than merely unfashionable.
Watch what happens next, because it is the clearest small example of absorption in the book. In June 2025 the newsletter runs a headline called Context Engineering: Much More than Prompts, and a new term arrives: 0.00, 0.00, 2.06, 4.64 — and then 1.97, 1.07. It rises for a year and fades in one.
Both terms describe the same job: getting the right words in front of the model. What changed is who does it. Here is the definition that circulated the week the new term arrived, from Andrej Karpathy, quoted in the newsletter on 25 June 2025:
In every industrial-strength LLM app, context engineering is the delicate art and science of filling the context window with just the right information for the next step. Science because doing this right involves task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history, compacting … Too little or of the wrong form and the LLM doesn't have the right context for optimal performance. Too much or too irrelevant and the LLM costs might go up and performance might come down.
Andrej Karpathy, quoted in AI News, 2025-06-25
Read the list of ingredients. Retrieval, tools, state, history, compaction — every one of those is code that runs on every turn, not a sentence a person writes once. Another contributor in the same issue gave the mental model directly: “just as an operating system curates what fits into a CPU's RAM, we can think about context engineering as packaging and managing the context needed for an LLM to perform a task.”
That is the transition, described by the people making it. The discipline did not fail. It was promoted into the harness, and things inside the harness do not get discussed.
The vocabulary of a craft disappears when the craft becomes a subroutine.
Before the adoption curve, the thing itself, because it is simpler than its acronym suggests and the newsletter explained it on the day it appeared. On 25 November 2024:
The Model Context Protocol (MCP) is an open protocol that enables seamless integration between LLM applications and external data sources and tools. Similar to the Language Server Protocol, MCP standardises how to integrate additional context and tools into the ecosystem of AI applications.
AI News lede, 2024-11-25
The Language Server Protocol comparison is the whole idea, and for anyone who has written software it is the only explanation needed. Before LSP, supporting n languages in m editors meant writing n × m integrations: a Python plugin for VS Code, a different Python plugin for Vim, and so on for every pair. LSP made it n + m. Each language ships one server that speaks the protocol; each editor ships one client; anything works with anything.
MCP does that for models and the things they need to reach. Before it, connecting a model to your database, your issue tracker and your filesystem meant bespoke glue for each model-tool pair, rewritten whenever either end changed. After it, a tool ships one server and any model that speaks MCP can use it.
The same issue lays out the parts, and they are deliberately unexciting:
Resources: any kind of data that an MCP server wants to make available to clients. This can include file contents, database records, API responses, live system data, screenshots and images, log files, and more. Each resource is identified by a unique URI and can contain either text or binary data.
AI News lede, 2024-11-25
A server exposes resources and tools. A client — the model's application — connects and consumes them. Resources are addressed by URI. There is nothing in that sentence a web engineer from 2005 would find strange, and that is the point: MCP is not a clever idea about intelligence, it is an integration standard, and integration standards win or fail on adoption rather than on cleverness.
What it looks like once it works is a list of company names. By September 2025:
Mistral Le Chat adds 20+ MCP connectors and “Memories.” Le Chat now plugs into Stripe, GitHub, Atlassian, Linear, Notion, Snowflake, and more, with fine-grained access controls and persistent, user-editable memory. This turns Le Chat into a single surface for cross-SaaS action and retrieval, while remaining enterprise-manageable.
AI News, Twitter recap, 2025-09-02
That is the n + m promise cashed: one client, twenty servers somebody else wrote, and the interesting engineering has moved to access control. Which is where the rest of this chapter comes in, because none of it happened on launch day.
On 25 November 2024, Anthropic published the Model Context Protocol — a specification for how a model talks to external tools. Here is the newsletter's opening line that day, in full:
AI News, 2024-11-25
claude_desktop_config.jsonis all you need.
The coverage underneath it is careful and technically accurate — it walks through resources, prompts, tools, transports and sampling, and notes that the docs make solid recommendations on security. Then it reports the reception:
The launch partners Zed, Sourcegraph, and Replit all reviewed it favorably, however others were a bit more critical or confused. Hacker News is already recalling XKCD 927.
AI News, 2024-11-25
XKCD 927 is the comic about competing standards, in which an attempt to unify fourteen of them produces fifteen. That was the informed reaction on day one, from people who had read the spec. Community-surface density for the month: 2.3.
Nothing much happens for two months. Then it climbs: 3.6, 13.3, 22.7, and a peak of 38.8 in March 2025 — the month a competitor adopted it.
Four months from publication to peak, with the inflection at somebody else's decision. That shape is not a failure of the launch; it is what a protocol's adoption curve has to look like. A protocol is worth exactly nothing until a second party implements it, so the interesting event is never the release. It is the first adoption you did not control.
Then the familiar decline: 25.3, 19.5, 14.3, down to 6.6 by early 2026. Not because MCP failed — by then the archive records adoptions by most of the major vendors it covers — but because it stopped being worth mentioning. A protocol everyone implements generates no more argument than a file format.
So the launch date told you almost nothing, and the informed day-one reaction told you less. The event that mattered was four months later, and it was somebody else's decision.
On 13 June 2025 the newsletter's headline is Cognition vs Anthropic: Don't Build Multi-Agents / How to Build Multi-Agents. Two well-resourced companies published directly contradictory architecture guidance close enough together that a daily newsletter covered them in one line.
That is worth pausing on, because it is what an engineering discipline looks like before it has patterns. The orchestration line rising twenty-threefold is not a field converging on how to build agents. It is a field arguing about it in public at increasing volume, and the archive's headlines make the disagreement legible in a way the density series cannot: Every 7 Months: The Moore's Law for Agent Autonomy in March 2025, Claude Agent Skills — glorified AGENTS.md? or MCP killer? in October, Agentic Engineering: WTF Happened in December 2025? in February 2026. Three different framings of the same unsettled question, ten months apart.
Abstract argument aside, the harness had products, and the archive keeps a running tally of which ones people were talking about. Counting the named tools across both text surfaces:
| Tool | 24H1 | 24H2 | 25H1 | 25H2 | 26H1 | 26H2 |
|---|---|---|---|---|---|---|
| Claude Code | 0.0 | 0.0 | 2.7 | 5.3 | 12.9 | 6.6 |
| Codex | 0.3 | 0.0 | 1.1 | 2.9 | 8.5 | 6.9 |
| Cursor | 0.0 | 0.8 | 1.7 | 1.7 | 3.3 | 3.4 |
| Copilot | 4.0 | 0.9 | 1.4 | 1.6 | 1.9 | 1.0 |
| Windsurf | 0.0 | 0.2 | 1.7 | 0.7 | 0.3 | 0.0 |
| Aider | 0.3 | 0.9 | 1.6 | 0.6 | 0.1 | 0.0 |
Four different shapes in one table. Copilot is the category in early 2024 — at 4.0 it is discussed more than every other coding tool in this table combined — and it never rises again; by the end it is a footnote in its own market. Windsurf and Aider both climb to a respectable 1.6 or 1.7 in the first half of 2025 and then go to zero, which in this archive is what acquisition and abandonment look like from outside; the series cannot tell you which, only that people stopped saying the name.
Claude Code and Codex do the thing nothing else in this book does: they go from literally absent — 0.0 in both halves of 2024 — to the two most discussed tools in the corpus inside eighteen months. Neither existed as a name when the harness argument started. Both were shipped into the argument while it was running.
And the archive is honest about how badly it can measure this, in a lede from June 2025 that is worth the whole chapter:
Anj from the newly rebranded a16z points out that there is a way to track background coding agent PRs in open source, and it's not much of a surprise that OpenAI Codex has something like 91.9% market share — but these numbers don't capture Claude Code's contributions, and Cursor's Background Agents are still prelaunch.
AI News lede, 2025-06-20
A precise number, a named source, and in the same breath the two reasons it is wrong. The denominator is public pull requests, which is where one product works and the others do not. Six months later Claude Code is the most-discussed tool in the archive. The 91.9% was never a lie; it was an answer to a question nobody had checked was the question they meant.
Market share of the thing you can count is not market share.
There is one line in this chapter's data that is larger than everything else and that I have been saving.
Evaluation language — eval, evaluation, benchmark —
runs at 38.37 per ten thousand words of announcement space in early 2024. At the end of the
corpus it is 67.93. In practice space, 34.90 to 60.35. It
rises by about three quarters in both surfaces — the same slope on each, which is what a
field-wide shift looks like — and ends higher than any other term in this chapter by a factor
of two.
Resist the temptation to pool. Broadening the pattern to include
leaderboard and counting both text surfaces together makes this look far more
dramatic than it is — I had it at 49.8 rising to 81.0 before splitting it, which overstates
both the level and the drama. Split, it is 45.6 → 72.8 in announcement space and 43.3 → 65.7 in
practice: a real, moderate, field-wide rise of about half again, with practice space dipping in
the middle. Every large number in this book that was not split by surface turned out to be
smaller than it looked.
What that rise is made of changes, though, and the archive is candid about it. By mid-2025 the complaint is no longer that benchmarks are hard, but that scores have stopped carrying information:
Post-training and RL saturation: lateinteraction critiques benchmark contamination and the overemphasis on RL-tuned gains, warning that much of the recent apparent progress may be due to prompt/template alignment rather than general capability.
AI News, Twitter recap, 2025-06-02
That is a precise and uncomfortable claim: that a model tuned until it matches the exact phrasing a benchmark uses will score better without being better at anything. It is the same failure as contamination — the test leaking into the training — arriving by a subtler route, and it is why the second half of this chapter is about which benchmarks resist it.
One caution on measuring the worry itself. Contamination is a word other
industries use, and a pattern that counts it will happily count a thread about pesticides and
heavy metals in the food supply, which is exactly what mine did on 1 July 2025. The
contamination series in this chapter is small enough — never above 3.2 per ten thousand words —
that a handful of such false positives moves it. Read it as evidence that the anxiety exists,
not as a measurement of how large it is.
The reason is mechanical. When the model was the product, you compared models. When the model is a component inside a system you wrote — with a retry policy, a context strategy, a tool set and a stopping rule, every one of which you chose — there is no way to know whether any change helped except to measure it. The harness turned every team into a team that needs an eval suite.
Whether those measurements were any good is a separate and harder question. But the demand for them is not in doubt: it is the largest signal in this chapter and one of the largest in the corpus.
Track the layer, not the component. Between 2024 and 2026 the interesting
engineering moved from choosing a model to building the thing around it, and every series in
this chapter says so at once.
A protocol's launch date is not its adoption date. Watch for the first
implementation you did not control; that is the event.
Re-derive what your metrics mean, on a schedule. harness went
from meaning an eval runner to meaning an agent loop while its count rose eighteenfold, and no
automated check anywhere would have caught it.
Budget for the eval suite. Once the model is a component in a system you
wrote, measuring the system is the only way to know whether a change helped.