Chapter 15
Why did none of this get caught by better statistics?
Here is a question with one obviously correct answer: by how much did this field's interest in fine-tuning decline between early 2024 and the middle of 2026?
Below are six answers. They are computed from the same archive, over the same window, with no arithmetic errors and no cherry-picking — these are simply the six most natural ways to ask the question.
| How you ask | One row is | 2024H1 | 2026H2 | Change |
|---|---|---|---|---|
| Raw mentions per issue | an issue | 84.4 | 3.1 | −96% |
| Mentions per 10⁴ words, announcement space | a word of the Twitter recap | 34.9 | 1.9 | −94% |
| Mentions per 10⁴ words, whole issue | a word of the issue | 33.0 | 5.3 | −84% |
| Mentions per 10⁴ words, practice space | a word of the Reddit recap | 19.4 | 7.0 | −64% |
| Share of issues mentioning it at all | an issue | 100% | 65% | −35% |
| Messages per month, busiest fine-tuning community | a Discord message | 7,208 | 19,400 | +169% |
The range is 265 percentage points, and it includes a sign change. One of these units says the subject nearly vanished. Another says it grew by more than half again. Every one of them is defensible, and I would have trouble telling you that any is wrong.
Nothing statistical happened here. The spread was fixed before a single number was computed, by decisions nobody writes down.
Every measurement over a corpus silently answers four questions. They are not statistical questions and there is no procedure for them.
What is one subject? A model checkpoint, a model family, a lab, a
technique.
What is one document? An issue, a section of an issue, a paragraph, a
channel-day.
What counts as one observation? A string match, a tagged entity, a claim made
about the subject, a person mentioning it.
What is the denominator? Words, documents, days, people, messages — or
nothing, if you report raw counts.
Walk the table with those four in hand and every row differs from its neighbours on exactly one of them.
Rows one and five share a document (an issue) and differ on the observation: one counts every mention, the other counts whether there was at least one. That single change moves the answer from −96% to −35%, because a subject can be mentioned once in almost every issue while its share of the text collapses.
Rows two, three and four share everything except the document. Narrowing from “an issue” to “the part of the issue written about announcements” moves it to −94%; narrowing to “the part written about practitioners” moves it to −64%. The two populations were doing different things, and the whole-issue number is a blend of them weighted by how much of each the publication happened to carry that year.
Row six changes the subject entirely. It stops counting words about fine-tuning and starts counting messages in the busiest community where people fine-tune models. That is a different population — participants rather than coverage — and it is the only row that measures activity rather than attention. It is also the only row that goes up.
This is the part I found genuinely uncomfortable, because my instinct when a result looks shaky is to reach for a better method, and none of the available methods touch this.
| Tool | What it protects you from | What it says about the unit |
|---|---|---|
| Larger sample | Sampling noise | Nothing |
| Confidence interval | Overstating precision | Nothing |
| Significance test | Reading noise as signal | Nothing |
| Multiple-comparison correction | Fishing across many tests | Nothing |
| Censoring-aware estimator | Unfinished observations | Nothing |
| Change-point detection | Missing a break in a series | Nothing — it finds breaks within a unit |
Every tool in that list operates on a quantity that has already been defined. Give any of them a larger sample and they will describe the wrong quantity more precisely. Precision and validity are orthogonal, and almost all the tooling — and almost all the training — is on the precision side.
The significance test is the clearest case. Ask of any row in that table whether the change could be sampling noise, and the answer is no, overwhelmingly, for all six rows simultaneously — including the one that goes up while the others go down. A test that confirms every mutually contradictory answer is not adjudicating between them. It was never asking that question.
Statistics tells you how much to trust the number you computed. It has nothing to say about whether you computed the number you meant.
The quantity you intend to measure has a name — the estimand — and the useful discipline is writing it as a sentence that somebody could disagree with.
“Fine-tuning declined” is not an estimand. It contains no subject definition, no population, no denominator and no window. Compare:
The share of text in this newsletter's Twitter-derived recap devoted to fine-tuning, measured as regex matches per 10,000 words of that section, fell 94% between the first half of 2024 and the second half of 2026.
That sentence is long, and it is the point. Every clause in it is a decision that could have gone another way, and stating them makes the disagreement possible. Someone can now tell you that the Twitter recap is the wrong population for the decision you are making, which is a useful thing to be told and impossible to say about “fine-tuning declined”.
Most arguments about data are arguments about estimands, conducted as though they were arguments about method. Two people looking at the same corpus and reaching opposite conclusions are usually both computing correctly.
There is no universally right unit, which is why this cannot be automated. There is a rule that resolves nearly every case in practice: the unit should match the population your decision affects.
| The decision | The population it affects | The unit to measure |
|---|---|---|
| Should we publish a post about our fine-tuning work? | People who read announcements | Density in announcement coverage |
| Should we keep the fine-tuning path in the product? | People who would use it | Activity in practitioner communities |
| Should we hire someone to maintain the tuning pipeline? | Our own engineers | Neither — internal usage of the pipeline |
The third row matters as much as the other two. Sometimes the right answer is that the corpus in front of you does not contain the population you care about, and the correct move is to stop measuring it rather than to measure it more carefully.
The whole of this chapter compresses into one fill-in-the-blank, and it takes about thirty seconds:
One row is ______, drawn from ______, and I am counting ______ per ______, over the window ______, in order to decide ______.
If any blank is hard to fill, the number does not mean anything yet. If the last blank is empty, you are measuring for the pleasure of it, which is allowed but should be admitted. And if two people fill the blanks differently, they will get different answers and neither will be wrong — that is the moment to argue about the sentence rather than about the result.
None of this is hard. It is just early, and everything that makes analysis feel rigorous — the tests, the intervals, the diagnostics, the plots — happens later, which makes it very easy to do all of it beautifully to a quantity nobody meant to measure.