Chapter 14
You are about to build on a model. How long will it stay relevant?
The model tag gpt-4o-2024-08-06 appears in this archive in four
issues, across two days. The tag gemini appears across
973. In August 2024 both were things you could have written into a config file
and shipped.
That spread is the practical problem this chapter is about. Every integration built on a model is a dependency with an expiry date, the date is not published, and the difference between the shortest and longest lives in this corpus is a factor of nearly five hundred. So: given a model you are about to build on, what is the distribution?
The obvious approach — take every model, subtract its first mention from its last, average — is wrong in a way that gets worse the more recent your data is. A model first mentioned last month and still being discussed today has not had a one-month life. Its life is unfinished, and counting it as one month drags the average down. Worse, it drags it down hardest for the newest models, which are exactly the ones you want to compare against the old ones.
This is the standard problem in medical statistics — patients still alive when the study ends — and the standard fix is the Kaplan-Meier estimator. Rather than averaging lifespans, it walks forward in time and asks, at each moment when something died, what fraction of the things still being observed died then. Multiply those fractions together and you get a curve: the probability of surviving past any given age. Things still alive at the end contribute to the denominator for as long as they were watched and then drop out without counting as deaths. That is called right-censoring, and handling it is the entire reason to use this method.
Here, birth is a model's first mention, and death is its last mention, provided that last mention is more than 90 days before the archive ends. Models mentioned recently are treated as alive.
Run that over every model tag appearing in at least three issues — 255 of them — and the answer to “how long will it stay relevant” is a curve rather than a number.
The median is 137 days. Four and a half months from a model's first mention to the halfway point of its time in the conversation.
Stated as a decision table, it is blunter:
| If you build on it today | Still discussed — checkpoint | — family |
|---|---|---|
| 1 month | 84% | 87% |
| 3 months | 62% | 77% |
| 6 months | 41% | 60% |
| 1 year | 21% | 45% |
| 18 months | 12% | 30% |
| 2 years | 5% | 21% |
If you pin an integration to a specific checkpoint today, there is a one in five chance anyone is still talking about it in a year, and a one in twenty chance in two. Pin to the family instead and those become roughly one in two and one in five.
Split the same 255 tags by cohort and something jumps out.
| Cohort | Tags | Median | Families | Median |
|---|---|---|---|---|
| All models | 255 | 137 d | 140 | 254 d |
| US frontier labs | 161 | 175 d | 49 | 315 d |
| Open-weights families | 100 | 117 d | 43 | 273 d |
| Chinese labs | 49 | 85 d | 21 | 398 d |
Read the left half of that table. Models from Chinese labs have a median life of 85 days against 175 for US frontier labs — less than half the shelf life. That is a clean, quotable, decision-relevant finding, and I published it, with a note that it might be a naming artifact.
It is a naming artifact. Collapse version and size suffixes so that
qwen3.5-235b-a22b, qwen3.6 and qwen3.8-max count as one
subject rather than three, and read the right half of the table. The cohort with the shortest
individual lives has the longest family lives of any group measured: 398 days
against 315 for US frontier labs.
The ranking inverts completely, on the same data, from one decision about what counts as a thing.
Two measurable causes, and neither is flattering to the original result.
Naming granularity. Inside this cohort, Chinese labs produce
5.44 distinct model tags per family against 4.03 for US frontier labs. More
named checkpoints per underlying thing means each individual name is shorter-lived by
construction, before anything about quality or adoption enters. A lab that ships
kimi-k2, kimi-k2-0905, kimi-k2-thinking and
kimi-k2-turbo-preview will look, at tag level, like four models that each died
young.
Cohort age. The median Chinese tag first appears on 2025-07-29. The median US frontier tag first appears on 2024-09-13 — eleven months earlier. The estimator handles censoring correctly, but a cohort concentrated in the final year of the corpus is estimated from far less evidence than one spread across all three, and its curve is correspondingly jumpy: the Chinese tag curve below sits flat at 12.9% for five straight months because almost nothing in that group is old enough to produce another event.
The corrected figure rests on 21 Chinese families. It is the better estimate — it is measuring subjects rather than strings — and it is not a precise one. Treat “longest-lived” as “not shortest-lived, and probably above average”, and do not build a procurement policy on the difference between 398 and 315.
| Cohort | Tags per family | Median first appearance | Median observed span |
|---|---|---|---|
| Chinese labs | 5.44 | 2025-07-29 | 77 d |
| US frontier labs | 4.03 | 2024-09-13 | 138 d |
A median is an abstraction. What it describes, for anyone who has built on a model, is the specific experience of a successor arriving and not being a drop-in. The archive records these as they happen, and they are rarely announced as regressions.
July 2024, on the model that replaced the one everyone had standardised on:
A vibrant discussion compared GPT-4o and Llama 405 tokenizers, highlighting GPT-4o's regression in coding language token efficiency versus its predecessor, GPT-4t — GPT-4o yielding more tokens in XML than GPT-4t, signaling a step back in specialized tokenizer performance.
AI News, Discord recap, 2024-07-17
Nothing about that is a failure of the newer model, which was better at most things and cheaper. It is a failure of the assumption that better is a scalar. If your costs are denominated in tokens and your payloads are XML, the upgrade is a price rise.
Two years later, the same shape at the frontier, and this time with a number on it:
CraigVG highlights a significant regression in long-context retrieval performance between Opus 4.6 and 4.7, with MRCR v2 scores dropping from 78.3% to 32.2%. However, Boris explains that MRCR is being phased out in favor of Graphwalks, which better reflects real-world usage and applied reasoning over long contexts.
AI News, Reddit recap, 2026-04-16
Read the second half of that carefully, because it is the more interesting half. A point release drops forty-six points on a published benchmark, and the response is not that the regression is disputed but that the benchmark is being retired. Both things can be true. But it means the instrument you would have used to detect the regression is itself on the same clock as the model, and retiring first.
The dependency you are managing is not the model. It is the pair: the model, and the measurement you trusted to tell you it still worked.
One thing this number is not: a switch-off date. GPT-4 remained callable long
after the conversation moved on, and a model dropping out of this archive says nothing about
whether its endpoint still answers.
What it does measure is how long the field keeps paying attention, and that turns out to be the clock that governs most of what actually goes wrong with an integration:
Pin to a family, not a checkpoint, wherever the API allows it. The whole
difference between the two columns of that decision table is this one choice.
Expect checkpoint churn on roughly this cadence for anything keyed to a
specific checkpoint. Note what the 137 days actually measures: how long a checkpoint keeps
being talked about, which is a hypothesis about how long it keeps being supported,
not a measurement of it. Vendors deprecate on their own schedule, and the two clocks are
correlated rather than identical — so use this to decide what to keep loosely coupled, not to
set a maintenance calendar.
Depreciate checkpoint-specific work on the same clock. Prompts tuned to one
model's quirks, few-shot examples chosen for its failure modes, and thresholds calibrated to its
scores are all assets with a four-month half-life. Anything you would be unwilling to redo twice
a year should not depend on a specific checkpoint.
Choose the ecosystem, not the score. What persists across a family is the
tooling around it — the quantizations, the serving stack, the community's accumulated tricks.
That is what you are really adopting, and it outlives any individual set of weights by a factor
of about two.
A model is a dependency whose version you do not control, on a release cadence nobody publishes, with a median attention span of four and a half months.
That is a worse deal than almost any library you have ever taken, and it is priced into nothing. The single cheapest mitigation is the one in the table above: depend on the family name, and let the lab decide which weights are behind it.