Written ForwardsChapter 14

Chapter 14

The half-life of a dependency

You are about to build on a model. How long will it stay relevant?

The model tag gpt-4o-2024-08-06 appears in this archive in four issues, across two days. The tag gemini appears across 973. In August 2024 both were things you could have written into a config file and shipped.

That spread is the practical problem this chapter is about. Every integration built on a model is a dependency with an expiry date, the date is not published, and the difference between the shortest and longest lives in this corpus is a factor of nearly five hundred. So: given a model you are about to build on, what is the distribution?

Why you cannot just average the lifespans

The obvious approach — take every model, subtract its first mention from its last, average — is wrong in a way that gets worse the more recent your data is. A model first mentioned last month and still being discussed today has not had a one-month life. Its life is unfinished, and counting it as one month drags the average down. Worse, it drags it down hardest for the newest models, which are exactly the ones you want to compare against the old ones.

Survival analysis, in one box

This is the standard problem in medical statistics — patients still alive when the study ends — and the standard fix is the Kaplan-Meier estimator. Rather than averaging lifespans, it walks forward in time and asks, at each moment when something died, what fraction of the things still being observed died then. Multiply those fractions together and you get a curve: the probability of surviving past any given age. Things still alive at the end contribute to the denominator for as long as they were watched and then drop out without counting as deaths. That is called right-censoring, and handling it is the entire reason to use this method.

Here, birth is a model's first mention, and death is its last mention, provided that last mention is more than 90 days before the archive ends. Models mentioned recently are treated as alive.

Run that over every model tag appearing in at least three issues — 255 of them — and the answer to “how long will it stay relevant” is a curve rather than a number.

Figure 23 · How long a model stays in the conversationKaplan-Meier curves over 255 model tags and the 140 families they collapse into. Death is 90 days of silence before the archive ends; anything mentioned more recently is treated as alive and censored.
0%25%50%75%100%0 mo6 mo12 mo18 mo24 mo137 d254 dall familiesall model tagsstill discussed

The median is 137 days. Four and a half months from a model's first mention to the halfway point of its time in the conversation.

Stated as a decision table, it is blunter:

Table 11 · Probability a model is still being discussed after…
If you build on it todayStill discussed — checkpoint— family
1 month84%87%
3 months62%77%
6 months41%60%
1 year21%45%
18 months12%30%
2 years5%21%

If you pin an integration to a specific checkpoint today, there is a one in five chance anyone is still talking about it in a year, and a one in twenty chance in two. Pin to the family instead and those become roughly one in two and one in five.

The result I published, which was wrong

Split the same 255 tags by cohort and something jumps out.

Table 12 · Median survival at tag level and at family level
CohortTagsMedianFamiliesMedian
All models255137 d140254 d
US frontier labs161175 d49315 d
Open-weights families100117 d43273 d
Chinese labs4985 d21398 d

Read the left half of that table. Models from Chinese labs have a median life of 85 days against 175 for US frontier labs — less than half the shelf life. That is a clean, quotable, decision-relevant finding, and I published it, with a note that it might be a naming artifact.

It is a naming artifact. Collapse version and size suffixes so that qwen3.5-235b-a22b, qwen3.6 and qwen3.8-max count as one subject rather than three, and read the right half of the table. The cohort with the shortest individual lives has the longest family lives of any group measured: 398 days against 315 for US frontier labs.

The ranking inverts completely, on the same data, from one decision about what counts as a thing.

Figure 24 · The same models, counted two waysThe lower curve counts each named checkpoint as a separate subject; the upper one collapses version and size suffixes into families. The long flat stretches in the lower curve are periods with too few subjects old enough to produce an event.
0%25%50%75%100%0 mo6 mo12 mo18 mo24 mo85 d398 dChinese familiesChinese model tagsstill discussed

Why it inverts

Two measurable causes, and neither is flattering to the original result.

Naming granularity. Inside this cohort, Chinese labs produce 5.44 distinct model tags per family against 4.03 for US frontier labs. More named checkpoints per underlying thing means each individual name is shorter-lived by construction, before anything about quality or adoption enters. A lab that ships kimi-k2, kimi-k2-0905, kimi-k2-thinking and kimi-k2-turbo-preview will look, at tag level, like four models that each died young.

Cohort age. The median Chinese tag first appears on 2025-07-29. The median US frontier tag first appears on 2024-09-13 — eleven months earlier. The estimator handles censoring correctly, but a cohort concentrated in the final year of the corpus is estimated from far less evidence than one spread across all three, and its curve is correspondingly jumpy: the Chinese tag curve below sits flat at 12.9% for five straight months because almost nothing in that group is old enough to produce another event.

The corrected figure rests on 21 Chinese families. It is the better estimate — it is measuring subjects rather than strings — and it is not a precise one. Treat “longest-lived” as “not shortest-lived, and probably above average”, and do not build a procurement policy on the difference between 398 and 315.

Table 13 · The two mechanisms behind the reversal
CohortTags per familyMedian first appearanceMedian observed span
Chinese labs5.442025-07-2977 d
US frontier labs4.032024-09-13138 d

What churn feels like from underneath

A median is an abstraction. What it describes, for anyone who has built on a model, is the specific experience of a successor arriving and not being a drop-in. The archive records these as they happen, and they are rarely announced as regressions.

July 2024, on the model that replaced the one everyone had standardised on:

A vibrant discussion compared GPT-4o and Llama 405 tokenizers, highlighting GPT-4o's regression in coding language token efficiency versus its predecessor, GPT-4t — GPT-4o yielding more tokens in XML than GPT-4t, signaling a step back in specialized tokenizer performance.

AI News, Discord recap, 2024-07-17

Nothing about that is a failure of the newer model, which was better at most things and cheaper. It is a failure of the assumption that better is a scalar. If your costs are denominated in tokens and your payloads are XML, the upgrade is a price rise.

Two years later, the same shape at the frontier, and this time with a number on it:

CraigVG highlights a significant regression in long-context retrieval performance between Opus 4.6 and 4.7, with MRCR v2 scores dropping from 78.3% to 32.2%. However, Boris explains that MRCR is being phased out in favor of Graphwalks, which better reflects real-world usage and applied reasoning over long contexts.

AI News, Reddit recap, 2026-04-16

Read the second half of that carefully, because it is the more interesting half. A point release drops forty-six points on a published benchmark, and the response is not that the regression is disputed but that the benchmark is being retired. Both things can be true. But it means the instrument you would have used to detect the regression is itself on the same clock as the model, and retiring first.

The dependency you are managing is not the model. It is the pair: the model, and the measurement you trusted to tell you it still worked.

What this is a clock for

One thing this number is not: a switch-off date. GPT-4 remained callable long after the conversation moved on, and a model dropping out of this archive says nothing about whether its endpoint still answers.

What it does measure is how long the field keeps paying attention, and that turns out to be the clock that governs most of what actually goes wrong with an integration:

Four things to do with a 137-day median

Pin to a family, not a checkpoint, wherever the API allows it. The whole difference between the two columns of that decision table is this one choice.
Expect checkpoint churn on roughly this cadence for anything keyed to a specific checkpoint. Note what the 137 days actually measures: how long a checkpoint keeps being talked about, which is a hypothesis about how long it keeps being supported, not a measurement of it. Vendors deprecate on their own schedule, and the two clocks are correlated rather than identical — so use this to decide what to keep loosely coupled, not to set a maintenance calendar.
Depreciate checkpoint-specific work on the same clock. Prompts tuned to one model's quirks, few-shot examples chosen for its failure modes, and thresholds calibrated to its scores are all assets with a four-month half-life. Anything you would be unwilling to redo twice a year should not depend on a specific checkpoint.
Choose the ecosystem, not the score. What persists across a family is the tooling around it — the quantizations, the serving stack, the community's accumulated tricks. That is what you are really adopting, and it outlives any individual set of weights by a factor of about two.

A model is a dependency whose version you do not control, on a release cadence nobody publishes, with a median attention span of four and a half months.

That is a worse deal than almost any library you have ever taken, and it is priced into nothing. The single cheapest mitigation is the one in the table above: depend on the family name, and let the lab decide which weights are behind it.