Chapter 1
What is this thing, and why would anyone read three years of a newsletter?
On the evening of 6 December 2023, someone published a summary of what the internet had said about AI that day. It ran to a few thousand words, it was assembled mostly out of Discord logs, and it opened by apologising for itself.
Hi alpha testers! That's right, there's now a custom intro for these newsletters. We're very flattered that hundreds of you have somehow found this crappy MVP and so I decided to put in a little last-mile human touch commentary.
AI News, issue 1
The headline that day was a question — Is Google's Gemini… legit? Gemini had launched the previous afternoon to an enormous marketing push, and the issue's verdict was that the marketing was great and the headline benchmark number was suspicious, because the MMLU claim rested on 32-shot chain-of-thought prompting rather than the single-shot number everyone else reported. Then it says something that no retrospective would ever contain:
We will know more on Dec 13th.
It did not know. That is the entire value of the thing.
The newsletter is called AI News, and it is written and edited by Shawn Wang, who
signs everything swyx — the same person behind the Latent Space podcast and
the AI Engineer conferences, both of which the archive covers as news while he is running them.
It changed publisher twice in three years, which you can read off the sender address on the
emails: ainews@buttondown.email, then news@smol.ai, then
swyx+ainews@substack.com. It ran from that first issue to 6 August 2026, which is
where this book's copy of the archive stops — 690 issues, 15,265,094 words,
covering 82% of the weekdays in between. The median issue is about 24,000 words long, which is a
third of a short novel, published daily, about the previous day.
Naming him matters more than courtesy. This book is going to argue, repeatedly, that the archive is one editor's view with one editor's taste, and that caveat is worth nothing if the editor is anonymous. Everything here is a measurement of what one identifiable person, with a podcast and a conference and a position in the field he is reporting on, chose to point a scraper at. He appears in this book as “the editor” from here on, for readability, and it is always the same person.
Each issue reads the same set of places and summarises them separately: a Twitter recap built from a declared list of accounts, a Reddit recap from a declared list of subreddits, and — until March 2026 — a Discord recap assembled from message logs across dozens of servers.
A note on dates, since they will matter. Every date in this book is the day an issue covers, not the day it was sent — an issue about Monday goes out on Tuesday morning, and the archive stores both. The covered day is the one that matters for a chronology, so that is the one used throughout.
Each issue also tells you, every single day, exactly what it looked at:
AI News for 12/1/2025-12/2/2025. We checked 12 subreddits, 544 Twitters and 24 Discords (205 channels, and 9665 messages) for you. Estimated reading time saved (at 200wpm): 697 minutes.
AI News, standard header
That header is printed 690 times and it is the most useful line in the whole corpus: a source that declares its own sampling frame, daily, is a rare thing and it makes an enormous amount possible.
It is worth being exact about who wrote what, because the two halves of an issue are not the same kind of text. The recaps underneath the header are model-generated from the sampled sources. The passage above it — the headline, the opening claim, the judgement about what mattered that day — is written by a person, and across the whole archive it comes to 124,977 words, or 0.81% of the corpus. Everything else is summary.
You can watch the seam. The human passage regularly argues with the summaries below it, and it is the only place in the corpus where anyone commits to an opinion that could later look foolish.
It is not a steady artifact, and the shape of the unsteadiness is the first thing worth knowing about it. Issue length nearly quadrupled in the first six months as the Discord sampling widened, held around 25,000 words for two years, and then collapsed by a factor of five in the second quarter of 2026, when the newsletter dropped the Discord recap entirely. The cadence barely moved through all of it — roughly 65 issues a quarter, start to finish.
That collapse matters more than it looks. An issue in early 2024 is mostly a digest of chat logs; an issue in mid-2026 is mostly a digest of forum threads and tweets. They are not the same kind of document, and anything you count across both is partly counting the change.
Almost everything written about the last three years of AI was written backwards. It was written after DeepSeek R1, after the agent turn, after the Chinese open-weights labs took the lead — and so it is organised around the things that turned out to matter. That is what history is for, and it is also why it is nearly useless for the question this book asks, which is not what happened but how could you have known.
A daily archive is different in kind. It contains, in dated form:
History is written backwards, once the winners are known. This archive was written forwards, with the wrong guesses intact. That is the only reason it is worth reading.
Being honest about the limits early is not throat-clearing; it changes what the rest of the book is allowed to claim.
Every number in this book measures attention within one curated view of a field's public conversation. Not deployment. Not revenue. Not capability. When you read that fine-tuning fell by 95%, that means the newsletter's Twitter recap devoted 95% less of its text to fine-tuning — while over exactly that window the busiest community anywhere in the archive was a fine-tuning toolchain with 302,248 messages. Both facts are true. Neither one is “fine-tuning declined.”
The three surfaces this book compares have short names and longer, more accurate ones, and it is worth fixing both in mind now, because the short names do the work of memory and the long names do the work of truth:
Announcement space is lab-linked Twitter discourse: 544 accounts on one
editor's list, summarised by a model.
Community space is sampled Discord activity: message volume in 56
English-language servers, which measures how busy a room was, not what software anyone ran.
Practice space is local-model Reddit discourse: 12 subreddits weighted heavily
toward people running models on their own hardware.
A gap between them is evidence about where a subject was being discussed. It is a lead worth investigating, not a measurement of adoption, and nowhere in this book is it ground truth about what was deployed.
There is one more limit, and it is the sharpest, because it sets a floor under every number in this book. The recaps are written by a language model, and the model changed. The archive says which one, in a line under each heading — all recaps done by Claude 3 Opus, best of 4 runs — and across the corpus that line names eight different model families for the Discord recap alone, while 384 of 613 issues never declare the Twitter summarizer at all. Holding the recap heading fixed holds the source fixed; it does not hold the instrument fixed.
How much does that matter? Three days were published twice, the same news summarised by two different models — 13 May 2024, 18 July 2024, 6 August 2024. Comparing pattern densities between the two editions of one day varies the instrument while holding the news constant, and the answer is a median of 1.23×, with a maximum of 2.40×. So a fold-change of 1.2 is indistinguishable from a summarizer swap, and one of 5.5 is not. Wherever this book leans on a ratio, that is the noise floor it is leaning against.
There are three further limits worth stating plainly. The archive is one editor's view, with one editor's taste; it over-weights the English-language, US-and-China, open-weights-adjacent conversation and largely misses enterprise procurement, academic publishing outside the poster-on-Twitter tier, and everything happening in Chinese-language forums. Its sampling frame widened over time — 7 subreddits to 12, 384 Twitter accounts to 544 — which means any count of distinct things mentioned is partly counting the newsletter's own appetite. And it stops. Discord coverage ends in March 2026; the copy of the archive behind this book ends in August 2026. Anything that looks like a decline after those dates is the instrument.
What is left, after all of that, is still remarkable: a dated, daily, self-describing record of what one fast-moving technical field paid attention to, across nearly a thousand consecutive days, written by people who did not know how it ended. Everything that follows is an attempt to read it carefully — starting with the fact that it is not one record but three, kept side by side, and that they disagree.