Chapter 9
Which capability went open, commoditised and absorbed inside three years?
For the first half of this archive, the most active practitioner community in it was not arguing about agents or reasoning. It was making pictures. In the first half of 2024 image-generation language runs at 148.3 mentions per ten thousand words of the Reddit recap — thirteen times its density in announcement space, and the highest figure any topic reaches on any surface anywhere in this corpus.
That number falls to 4.7 by the end, and the reasons are mostly about who the newsletter was listening to rather than what was happening; the arithmetic is taken apart in the interlude before this part. What is left once you set the arithmetic aside is a genuine and unusually complete story, because image generation is the one capability in this book that went through the whole cycle — open, commoditised, absorbed — inside three years.
Start with pure noise: a grid of random pixels. Train a network to remove a little of that noise, over and over, and to do it conditioned on a piece of text. Run it a few dozen steps and the noise resolves into an image that matches the description. That is diffusion, and the counter-intuitive part is that the model never learns to draw. It learns to denoise, and drawing is what repeated denoising looks like from outside.
Doing this on raw pixels is expensive, so production systems work in a compressed latent space — a smaller representation an autoencoder can expand back into a picture — which is where the "latent" in Stable Diffusion's name comes from. The archive tracks the architecture being reworked underneath that idea throughout 2024:
a hierarchical pure transformer backbone for image generation with diffusion models that scales to high resolutions more efficiently than previous transformer-based backbones. Instead of treating images the same regardless of resolution, this architecture adapts to the target resolution, processing local phenomena locally at high resolutions and separately processing global phenomena in low-resolution parts of the hierarchy.
AI News lede, 2024-01-23
Image generation reached practitioners before text generation did, and it reached them with weights. Stable Diffusion ran on a gaming card, and the ecosystem that grew around it — fine-tunes, LoRAs, samplers, upscalers, interfaces — was the first place where large numbers of people built things with a model they had downloaded. That is why practice space in early 2024 is so dense with it, and it is also why the vocabulary of the fine-tuning chapter appears here first: LoRA was an image-model technique that text models borrowed.
Then the named products turn over, fast:
| Announcement space | 24H1 | 24H2 | 25H1 | 25H2 | 26H1 | 26H2 |
|---|---|---|---|---|---|---|
| Stable Diffusion / SDXL | 1.7 | 1.5 | 0.2 | 0.1 | 0.1 | 0.0 |
| DALL-E | 1.5 | 0.2 | 0.2 | 0.0 | 0.0 | 0.0 |
| FLUX | 0.0 | 2.2 | 1.1 | 1.7 | 0.6 | 2.1 |
| Midjourney | 0.2 | 0.9 | 0.7 | 0.4 | 1.3 | 0.0 |
| Imagen / Nano Banana | 2.8 | 0.2 | 1.0 | 4.3 | 0.8 | 0.4 |
Stable Diffusion and DALL-E, the two names that defined the category, are effectively gone from the conversation by 2025. FLUX — from a team of ex-Stability researchers — takes the open-weights lineage over almost immediately. This is the fastest turnover of leading products anywhere in the corpus, and it happened while the underlying capability was steadily improving, which is worth holding onto: a name disappearing here means a company lost, not that the technique did.
What that density is actually made of is worth seeing, because it is not people discussing image generation as a topic. It is people using it:
In /r/StableDiffusion, users share tips and tricks such as base prompts for realistic SDXL renders, colouring in with AI, and creating custom Stardew Valley player portraits.
AI News, Reddit recap, 2024-04-01
Regional prompting experiments with A8R8, Forge, and a forked Forge Couple extension allow more granular control over image generation, with the new interface supporting dynamic attention regions, mask painting, and prompt weighting to minimize leakage.
AI News, Reddit recap, 2024-04-02
Two days apart, and between them the whole character of the era: game portraits and a forked extension for controlling which prompt applies to which region of the canvas. This is a craft community with tooling, not an audience for launches — which is exactly why its density was so high, and exactly why it does not survive the sample widening.
The structural change arrives in 2025, and the archive catches it being described as a distinction rather than an upgrade:
the team's progress in creating native image generation for Gemini 2, highlighting its difference from text-to-image models.
AI News, Twitter recap, 2025-03-14
The difference is the whole chapter. A text-to-image model is a separate system you send a prompt to. Native generation means the same model that holds the conversation emits the image, so it can use everything it already knows — what you asked for three turns ago, the document you pasted, the correction you just made. Editing becomes a conversation instead of a new prompt, and the model can read images as fluently as it writes them.
Once that lands, a standalone image generator stops being a frontier product and becomes a feature of a chat assistant. Which is exactly what the density series shows, and exactly what it cannot explain on its own.
The most interesting afterlife is that the method outlived its application. Counted
on its own, diffusion in announcement space runs 4.5, 3.5, 7.5,
5.5, 4.1, 1.5 — peaking in the first half of 2025, well after image generation stopped being the
story. The reason is that researchers began applying denoising to text, producing language
models that generate a whole sequence in parallel and refine it, rather than one token at a
time. The archive covers LLaDA, a large language diffusion model, in February 2025.
The image-generation era ended by winning twice: its interface was absorbed into chat, and its mathematics was borrowed by text.
Video followed the same arc a year later and faster — Sora, Veo and Kling arrive as separate products in 2024 and 2025, and video-generation language peaks in the second half of 2024 at 12.1 before settling. By the end of the corpus the frontier releases are multimodal by default: one model that reads and writes text, images, audio and video, and the separate categories this chapter is named after have stopped being categories at all.
Read the four rows of that figure in the order they appear and you have the whole arc without needing the prose. Practitioners got there first, loudly, because the weights were downloadable and the hardware was a gaming card. The announcement layer arrived late, peaked in 2025 and declined gently, because a capability that has become a feature of a chat assistant is nobody's launch. Community space barely moves at all. And the editor's line — the only human one — ends higher than it started, which is the tell that nothing about the subject actually died.
Image generation is the completed example the rest of the book only has fragments of: open first, commoditised second, absorbed into a general model third, and its mathematics borrowed by the thing that replaced it.