Chapter 11
Which benchmarks are worth believing, and for how long?
In the last half-year of this archive, MMLU appears zero times in
announcement space. So does HumanEval. So do GSM8K,
MT-Bench and AIME. Every general-knowledge and fixed-problem-set
benchmark that defined how the field ranked models in 2024 reaches literal zero, in 46,815
words of text about model launches.
They were not discredited. There was no scandal and no replacement announcement. They stopped being mentioned, one at a time, over about two years — and the pattern of how long each one lasted turns out to be regular enough to plan around.
Take each benchmark's density inside the same fixed recap section, quarter by quarter. Find the quarter it peaks. Then find the first quarter after that where it has fallen below a quarter of its peak value — the point at which the field has functionally stopped using it.
| Benchmark | What it asks | Peak → below 25% of peak | Quarters |
|---|---|---|---|
| MT-Bench | rated answers | 2024Q2 → 2024Q3 | 1 |
| MMLU | multiple-choice recall | 2024Q2 → 2024Q4 | 2 |
| HumanEval | fixed coding problems | 2024Q2 → 2024Q4 | 2 |
| LMSYS Arena | human preference | 2025Q1 → 2025Q4 | 3 |
| FrontierMath | fixed maths problems | 2025Q4 → 2026Q3 | 3 |
| GPQA | graduate-level recall | 2025Q1 → 2026Q1 | 4 |
| AIME | fixed maths problems | 2025Q1 → 2026Q1 | 4 |
| SWE-bench | fix a real repository | 2026Q1 → still rising | — |
| ARC-AGI | solve unseen puzzles | 2026Q1 → still rising | — |
| Terminal-Bench | drive a shell to a goal | 2026Q3 → still rising | — |
The median is three quarters from peak to irrelevance. MT-Bench managed one. MMLU and HumanEval, the two benchmarks that appear in almost every model card of early 2024, managed two each. The most durable of the fixed-answer benchmarks, GPQA and AIME, managed four.
Three never fell at all, and they are the interesting ones. But first, why the others went.
A benchmark is a fixed set of questions with known answers. That design has a built-in expiry date, and it expires twice.
The first expiry is saturation. Once the leading models score 90% on something, the remaining 10% is disproportionately made of ambiguous items and mislabelled answers, and the benchmark can no longer rank anything. A test that everyone passes measures nothing, and the field's response is not to fix it but to stop citing it.
The second expiry is contamination, and it is worse because it is invisible. The questions are published so people can use them; being published means being on the public internet; being on the public internet means being in the next model's training data. At that point a high score may mean the model can reason, or may mean it has seen the answer, and nothing about the number distinguishes those.
You can watch the anxiety about this build. Contamination language in announcement space runs at 1.06 in early 2024, disappears entirely in the second half of that year, and returns at 1.92 in the final half-year — its highest value in the corpus, at the point where most of the benchmarks it refers to have already gone quiet.
By late 2025 it has stopped being a caveat and become a design requirement. Here is the newsletter describing a new benchmark in September 2025, and note where contamination sits in the list of features:
SWE-Bench Pro: a harder successor to SWE-Bench Verified with multi-file edits (avg ~107 LOC across ~4 files), contamination resistance (GPL/private repos), and tougher deps. Current top scores: GPT-5 = 23.3%, Claude Opus 4.1 = 22.7%, most others <15%.
AI News, 2025-09-22
Contamination resistance is listed between the edit size and the dependency difficulty, as an ordinary engineering property of the artifact. And the top score is 23.3%, which is what a benchmark looks like before it saturates.
If fixed questions saturate and leak, the obvious fix is to stop asking fixed questions. Show two anonymous answers to a human, ask which is better, and aggregate a few million of those into a ranking. That is the Chatbot Arena, and for a while it was the most-cited scoreboard in the field: 2.63 in early 2024, 8.32, and a peak of 12.04 in the first half of 2025.
Then it goes 6.93, 0.56, 0.43.
Two entries from April 2025 explain the fall better than the series does. The first is about what the ranking measures:
lm-sys highlighted the importance of style and model response tone on Arena, demonstrated in style control ranking … They updated their leaderboard policies to reinforce their commitment to fair, reproducible evaluations. [@vikhyatk] shared that this is the clearest evidence that no one should take these rankings seriously.
AI News, 2025-04-09
The operators found that response style and tone were driving the votes, and shipped a correction for it. A prominent practitioner read the existence of the correction as proof the whole thing was unsound. Both reactions are defensible, which is the problem.
The second is about what the ranking is a ranking of:
Llama 4 quietly dropped from 1417 to 1273 ELO, on par with DeepSeek v2.5.
AI News, 2025-04-14
A hundred and forty-four Elo points, because the checkpoint that had been evaluated was not the checkpoint that shipped. A score you cannot reproduce on the model you are able to call is not a score about that model.
Human preference is a real signal and also a measurement of formatting, length and tone. The moment it becomes the target, it measures the targeting.
Three benchmarks in this archive never fall below a quarter of their peak, and two of them are still climbing when the corpus ends. SWE-bench asks a model to fix a real issue in a real repository, and the test suite decides whether it worked. Terminal-Bench asks it to drive a shell to a goal state. ARC-AGI asks it to solve visual puzzles it has not seen, from a private holdout set.
What they have in common is not difficulty. It is that none of them has an answer key.
There is a verifier, not a lookup. The tests pass or they do not; the shell
reaches the goal state or it does not. Correctness is computed rather than looked up, which
makes this class of benchmark markedly harder to contaminate than one scored against a
published table of answers.
The task space is effectively unbounded. There are always more repositories and
more shell tasks. When the current set saturates you add harder instances, which is a data
problem rather than a redesign.
Partial credit is meaningful. A 23% pass rate leaves room to rank things.
A 94% multiple-choice score does not.
The same logic produced a structural fix for the quiz-style benchmarks that survived by changing what they are. From October 2025:
Rolling “Humanity's Last Exam”: CAIS released a dynamic fork of the well-known eval dataset that swaps easier questions for harder ones as models improve; gated to avoid contamination.
AI News, 2025-10-08
A benchmark that is a service rather than a file. It cannot saturate because it replaces its own easy items, and it cannot leak because the items are not all public. That is the same two defences the task benchmarks get for free.
Running enough real tasks is expensive, so the field did the other obvious thing: it asked a model to grade the output. LLM-as-judge language rises from 1.08 to 4.27, and by the end of the corpus it is denser than any individual benchmark except the task suites.
This solves cost and scale and does not solve validity. You have replaced a fixed set of questions whose answers might be in the training data with a grader whose own behaviour changes every time it is updated — and whose preferences, like the Arena's, include style. It is a reasonable tool for regression testing against yesterday's output. It is a weak instrument for deciding whether one model is better than another, and the archive does not contain a single result that resolves that.
Assume roughly three quarters of useful life from a public benchmark's peak.
Plan to replace your evaluation set on that cadence rather than standardising on one.
Prefer a verifier to an answer key. If correctness is computed — tests pass,
goal state reached, output parses and satisfies constraints — the benchmark resists both
saturation and leakage.
Distrust any score you cannot reproduce on the checkpoint you can actually call.
A hundred and forty-four Elo points separated the evaluated Llama 4 from the shipped one.
Keep a private held-out set drawn from your own traffic. It cannot leak, it
cannot saturate while your traffic changes, and it measures the only thing you actually care
about.
The field went through roughly a dozen scoreboards in thirty-two months. The ones it kept are the ones that ask a model to do something and then check whether it worked; the ones it dropped are the ones that ask a model a question someone already knows the answer to. That is a smaller claim than "benchmarks are broken", and a more useful one, because it tells you which kind to build.