April 24, 2026. The light comes up over Yuhang slowly and at a low angle, with a haze on the canal that the locals call yan and the meteorologists call particulate, both of them right in their own way. My alarm goes off at 6:15. I lie in bed for a few minutes because the alarm is, in practice, a suggestion, and then I get up and put water on for tea.
I live alone in Liangzhu, about 20 minutes by metro from the office. The kitchen window faces east. Today is hazy and I can see the building across the courtyard and not much else. I drink the tea standing at the counter because I have not yet committed to the day.
The kettle clicks off and I think, as I have thought every morning for the last six weeks, today is the day, and then, more honestly, today is one of the days it might be. The launch window has been moving for two months. We froze the model in the last week of March. Post-training was finalized on April 8. The technical report went through its final pass on the 19th, with three of us arguing about a single sentence for 80 minutes, until Su Yutong said just leave it and we left it. The Hugging Face repository has been staged since Sunday. We have been waiting, since Tuesday, for a decision none of us is responsible for and all of us are anticipating.
I take the metro at 7:20. I get off at 8:11 and walk to our building, past the noodle place that opens at 11, past the magnolia tree that has just finished flowering.
The twelfth floor at 8:15. Perhaps eleven people. Chen Wei is on the conference room couch where he has been sleeping at least three nights a week since February. I sit at my desk, wake the monitor. The progress bar on the upload terminal is at 97%. The bar is decorative. The actual upload finished overnight. What it is waiting for is the go-signal, which is what the entire twelfth floor is waiting for.
What is true about this floor, and what was true when I joined in 2023, is that there is no clock-in. No formal performance reviews. No KPIs. Most people walk out at six or seven. The reasoning is that nobody can sustain more than six to eight hours of high-quality output, so any system that demands more is a system that demands theatre. We are not performing the work. We are doing it.
At 8:40, Chen comes out of the conference room with his hair pressed flat on one side and his glasses smudged. He brings me tea without asking — a green tea from a small farm near Longjing Village that his cousin works at.
Out yet? he says.
No.
What are they doing.
I think they're waiting on the Huawei livestream.
The floor fills up toward nine. Nobody is talking very much. Everyone knows the model is launching today and nobody has anything in particular to say about it.
We finished the work two weeks ago. The launch is for other people.
I think about the 15 months that have brought us to this morning — the people who left, the Huawei engineers on the cots, Chen at the whiteboard at one in the morning — and I think about how those 15 months will, by lunchtime, become a single sentence in foreign press: DeepSeek released V4-Pro and V4-Flash on April 24, 2026. The sentence is true. Most of what is also true will not fit in the sentence.
At 9:20, the repository flips public. I refresh the page on my phone. The model card loads. I sit at my desk and feel, for a moment, nothing in particular. Then I feel something closest to relief but not relief, which I will later decide was probably just the feeling of a long thing ending.
I message Chen: 它出去了.
He replies in seven seconds: 吃过早饭没.
I have not. I get up to find something to eat.
January 20, 2025. R1 ships at 5 PM Beijing time. I am at my parents' apartment in Hefei because my mother has made it clear, by way of three separate WeChat messages, that early is the new on-time. My father is in the kitchen making jiaozi. The 春晚 gala is on in the living room, the volume up, the way it is every year whether anyone is watching or not. I am at the small desk in my old bedroom, which still has the desk lamp I used for homework and a shelf of textbooks nobody has thought to throw away.
The news arrives in waves over the following five days, and the waves are not the waves we expected. The first is professional — Hacker News, Reddit, Karpathy, LeCun. The second is financial. NVIDIA falls 17% on Monday, losing approximately $600 billion in market capitalization — the largest one-day decline by a single company in the history of the U.S. stock market. My father, who has not previously demonstrated any interest in my work, walks into the guest room and asks, in the polite voice he uses when trying not to sound impressed, what DeepSeek is. I tell him it is the company I work for. He nods and turns up the television, which my mother has tuned to CCTV-13. The anchor is saying our company name in the careful pronunciation reserved for words she has not said before. My father comes back 10 minutes later and asks if I would like a beer. We do not normally have a beer together. I say yes.
The third wave is at home. The founder's village in Zhanjiang hangs a red banner — warm congratulations on becoming the pride of your hometown. Hangzhou names DeepSeek one of the 杭州六小龙 — the Six Little Dragons. I learn we are one of the dragons from a SCMP article on the high-speed train back. I had not previously been aware that there was a list.
I return to the office early. There are 47 messages from headhunters in my inbox — 47 over 23 days. The packages are absurd. Two and a half times my pay. Three times. I do not reply to any of them. What I am thinking about is what is in the inboxes of the more senior people on the floor.
What was in their inboxes was larger. ByteDance, Tencent, Xiaomi, Alibaba — the offers for the eight or nine names on the R1 paper escalated through the spring. 100 million yuan a year. 80 million over four years. We could not match these numbers. DeepSeek had never raised outside capital. The options the lab had begun issuing were tied to a strike price nobody knew how to translate into an actual number. An option without a priced strike is not currency. It is a promise.
Then, on a Tuesday in February, Guo Daya told the team he was leaving.
Guo had been the architect of GRPO, first author on the R1 paper. He was 31, and indispensable in the sense that nobody had done his job before him and nobody had been trained to do it after him. He had been a quiet presence on the floor — soft-spoken, long-haired, the same gray hoodie for a week at a time — and he had been, in the small ways that someone is, a friend of mine. I had once had a long conversation with him about a paper of Hinton's we had both found unconvincing, and we had stayed in the small kitchen until past nine drinking tea and arguing.
His email was three sentences. The salary, when LatePost reported it, was rumored near 100 million yuan a year. The all-hands was the next morning. The founder was in the room, in the back, against the wall, but he did not speak. There were no questions, because the questions everyone wanted to ask were questions you do not ask in an all-hands. I sat through it and felt, briefly and shamefully, a small flicker of envy at a number I had not been offered, followed by irritation at Guo for timing his exit before V4 froze, followed by a residual affection that overwrote both and left me unsure of what I actually felt.
Su came to my desk that afternoon. She stood with one hand on the back of the empty chair and said, 他做了他的选择. 我们做我们的. He made his choice. We make ours. Then she went back to her desk.
By the end of March, the floor had lost five names the field knew. The people who left were not traitors for leaving. The calculus was different for each person, and was private, and was theirs. We came in. We went home at six. We came in again the next day.
What the outside world did not see, because the outside world was watching the departures, was that the lab was shipping. V3-0324 in March. R1-0528 in May — reduced hallucinations, function calling, the fixes that made R1 usable in production. V3.1 in August, which was the release where we made the decision that would define V4: thinking and non-thinking in a single model, not separate siblings. We had been building a standalone reasoning model internally — what would have been R2 — and the August release was the public signal that we had killed it and folded reasoning into the base architecture instead. V3.1-Terminus in September. Then V3.2-Exp later that month, which was the DeepSeek Sparse Attention rehearsal. Then, in December, V3.2 and V3.2-Speciale — which won gold at the IMO, CMO, ICPC World Finals, and IOI 2025. The week the Speciale results came in, two of the researchers whose names were on the olympiad papers had headhunter offers on their desks. The gold medals did not insulate against the market. Nothing did.
Through all of this, the work that occupied the second half of 2025 was also trying to train V4 on Huawei's chips. Beijing had communicated that the strategic preference was for V4 to train on Ascend — not just to serve, to train. The political logic was clear. The technical logic was what we lived inside.
Two engineers from Huawei's Shenzhen team arrived in October. One of them — I will call him Lin Zhao — moved into our small conference room within his first week. He was perhaps 35, Tsinghua electrical engineering, tall and thin with the hollow look of a person who has been on call too long. He smoked in the parking lot, and Chen began going down with him during debug sessions, because Chen had figured out that the fastest way past a stuck problem was a 15-minute walk away from the screen.
In early December, the founder convened a meeting — 15 people, two hours. Chen told me the next morning that the decision had been made to move V4 training back to NVIDIA hardware. Su told me, a week later, that the founder had said one sentence everyone remembered: 先把模型做出来. Let's make the model first.
The Huawei team pivoted to serving-side adaptation. By March, the Ascend 950 supernodes were matching our NVIDIA inference latencies. I saw Lin Zhao once more, in a video call. He had a haircut and looked, for the first time since I had met him, like a person who had slept.
We tried to train on Chinese chips and could not. We trained on NVIDIA chips. We served on Chinese chips, and in serving at scale we accomplished something real that was not what we had set out to accomplish. The distinction matters. It will, in most retellings, be lost.
By January 2026, the model was not yet trained. We had lost five senior researchers. The ones who remained were tired. We had three months.
By February 16 we had three months until the launch window and no model. The architecture that shipped on April 24 emerged from three long arguments between people who liked each other and who disagreed. A researcher I will call Jia won the first, about attention — she had been arguing since August that you needed to compress the key-value cache first and then run our sparse indexer on the compressed representation. The September V3.2-Exp release had been the public rehearsal for her design. The ablations in mid-February were very good: 27% of the compute, 10% of the memory. Jia took the next day off, which is the only time in 2.5 years I am aware of her doing so. Chen won the second, about the residual stream — his proof came in pieces over 11 days, a sketch on a whiteboard while Jia, in the room for a different reason, sat down and started asking questions. Su came through and stopped at the dashboard and looked at the loss curve for 40 seconds. Then she said, quietly, 这是地基. This is the foundation. The third argument, about post-training, was Su's, won over 14 months. I will come back to it.
By late February we had a 1.6 trillion parameter design and no confidence. A training run is not a launch. It is more like a long voyage on a ship that has been built mostly correctly but needs adjustment in dozens of small ways during the journey. The ship is mostly fine. The watches are mostly boring. The crew develops a relationship with the ship that is not romantic but is also not nothing.
The V4-Pro run started on February 28 at 03:14. The on-call engineer — a woman I will call Dai — watched the first 1,000 steps, sent a message to the cluster channel: 看起来正常 — looks normal — and went home. The run was scheduled for 42 days. It would take 46. The first two weeks were uneventful. The loss curve descended in the smooth, slightly noisy way that means the model is learning. I checked the dashboard on my phone in the evenings. Chen checked it more often. Su checked it less often, which I understood to be either confidence or fatalism.
The unplanned pause came at step 470,000, on the morning of March 18.
I was in the building. The dashboard showed a loss curve that had begun to climb. The slope was not steep. It was unmistakable. I called Chen, who had slept at home for the first time in 11 days.
Loss is climbing.
How fast?
Slow. About 0.02 per 1,000 steps for the last 20 minutes.
Is the indexer firing right?
I haven't checked.
Check.
Chen arrived at 8:47 in a taxi. Su arrived independently at 8:52 — she had seen the dashboard from her apartment and come in without anyone calling her.
The cluster bay is a long narrow room on the 11th floor, cast in the blue-green light of dashboards that are always on. Chen sat down next to me at the main terminal. Su stood behind us, watching the loss curve update every 30 seconds.
Show me the routing weights, Su said.
The histogram was wrong. Three of the 64 experts were receiving twice the token load they should have been. We pulled four checkpoints spanning the previous 8 hours. The clustering had begun 2,000 steps before the loss had noticed. The router had been drifting for hours.
Chen leaned back, quiet for 30 seconds. Then: The router is updating on features that are already stale by the time the update lands.
Su said, Anticipatory Routing. It was not a question. We had designed a scheme for exactly this failure mode — decoupling routing updates from the backbone so that during a spike, routing used historical parameters while features used current weights. The automatic threshold had been set too conservatively. The first 20 minutes of the climb had not triggered it.
Engage it manually, Su said.
Now? We haven't finished the diagnosis.
The diagnosis is the router. Engage it.
I looked at Chen. He nodded once. I engaged it at 12:14.
We sat in the cluster bay for the next four hours. Nobody ate lunch. Su did not sit down once. She stood with her arms crossed and watched the loss curve the way you watch a thing that is either going to turn around or is not. The climb slowed within 400 steps. It reversed within 900. By the time we left at four, the room smelled faintly of the green tea Chen had spilled at some point without either of us noticing.
The line in the technical report — the underlying principles remain insufficiently understood — refers to this episode. We argued about whether to print it. I thought it would be read as weakness. I was wrong. The line is in the paper.
The base model finished training on April 13: 训练完成. Training complete. No celebration. The base model was a base model. The model that would ship in 11 days was not yet built.
What followed was Su Yutong's domain. Her proposal: train five specialist models independently — math, competitive coding, agent use, instruction following, and a catch-all — each with its own reward model. Then merge them into a single student through a technique she called On-Policy Distillation. The student would learn to match all five specialists simultaneously, pulled toward each of them in distribution space.
The specialist training ran in parallel with the base model from late February. The math specialist, in its first three weeks, drifted into a regime where it was writing long chains of plausible reasoning with a single wrong step buried in the middle — the kind of step a human reader would nod past, because the notation was correct and only the content was wrong. We spent a week diagnosing what turned out to be a calibration problem in the verifier. It was the problem that most made me think about what these models are actually doing when they reason.
The merge began on April 5 and ran for nine days. Su was, during those nine days, the most focused I have ever seen another person be. She slept in the small conference room on a folding mattress. She took meals at her desk — rice in a plastic container she ate from without looking at, her eyes on the terminal. I would come in at eight and she would be in the same position she had been in when I left at seven, and I would not be certain she had moved. She was adjusting the merge weights by hand, twice a day, based on her reading of the rollout distributions. There was no automation for this. There was Su, and the terminal, and her judgment.
On the morning of April 14, the merge produced a model. Codeforces internal Elo: 3,206 — approximately the 23rd rank among human contestants. Putnam-2025: 12 out of 12. Chen ran the Codeforces evaluation five times. The number held. He sat at his desk for a long time and did not say anything.
Su wrote, in response, a single character: 嗯. Then she went back to work, because five more days remained before the model was ready to freeze.
The model froze on the 19th. On the 24th, at 09:20, the repository flipped public.
The week before V4 launched, the world arrived. April 16: Anthropic shipped Claude Opus 4.7, which was 9 points ahead of V4-Pro on SWE-Bench Pro. Chen came to my desk and said, 他们也在跑. They are also running. The lab got quieter for the rest of that day.
April 17: The Information reported that DeepSeek was in talks for its first external fundraising round — at least 2 billion yuan at a valuation above 70 billion. The company had refused outside capital for its entire existence. I learned about it from my cousin in Shanghai. The article noted the change in posture had been the result of a private decision in February that had not been explained to the broader team.
I had not been told. I sat with this at my desk for longer than I would have expected. The feeling was not anger. It was the recognition that I had spent 15 months doing work I believed in for an organization whose shape I understood, and the shape was changing, and the change had been decided without me. Which was appropriate — I was not senior enough to be in that room. Which was also the kind of thing you can know is appropriate and still feel the weight of.
Su came to my desk at four. 看到了. I saw it. 我也看到了. I also saw it. Six words. I have thought about them more often than I would have predicted.
April 23: OpenAI shipped GPT-5.5, which was 15 points ahead of us on agentic coding. Chen said at lunch, 他们的模型也是模型. Their model is also a model. He was not wrong. I did not find it particularly comforting.
Opus 4.7, the fundraising report, GPT-5.5 — we had hoped to ship into a quieter week. We were not going to.
The Hugging Face repository went public at 09:20. The Huawei livestream began at 09:47. Simon Willison published his post: DeepSeek V4 — almost on the frontier, a fraction of the price. We were almost on the frontier. On the dimensions where we were strongest, on it. We were close, and cheap, and open, and not first.
Su came to my desk near six and said, 休息一下, take a rest, and went home. The founder posted two characters in the team channel: 谢谢. Thank you. The message scrolled up along with everything else.
The launch announcement closed with a line from Xunzi:
不诱于誉,不恐于诽,率道而行,端然正己.
I read it on the metro home, on my phone. It is a line from a text written approximately 2,300 years ago. Whatever it meant in the announcement was what it meant.
I get off the train at 19:14. The locust trees by the metro entrance have just begun to flower. I take the long way home, past the small park with the koi pond. The koi are at the surface, mouthing at insects I cannot see. I sit on a bench for 10 minutes. The thing about a long project ending is that it does not feel like the thing the project has been building toward. It feels like an ordinary evening in which a pressure that has been present for a long time is now absent, and the absence is not immediately replaced by anything.
There is a noodle place at the corner of my street that stays open until 11. The owner is a woman from Shanxi who has, over four years, stopped asking me what I want. She nods at me and points at the corner table. A few minutes later she brings me biang biang noodles with chili oil and a pot of tea I did not order.
My phone buzzes. Hugging Face: DeepSeek-V4-Pro has crossed 10,000 downloads. I put the phone face-down. It buzzes again. A number I do not recognize. English. Hi, I'm reaching out from a major AI lab, and we have an opportunity I'd love to discuss. I read the first sentence and do not read the rest. I pour myself another cup of tea.
I think about the work still ahead. None of it is settled. The verdict on V4 will not be available for at least a year.
The path is the path.
I think about Guo Daya, somewhere in Beijing tonight doing good work. I feel about him the way you feel about people who used to sit near you and have moved on — fond, a little distant, not particularly resolved.
I finish the noodles. I walk home. Through the kitchen window the haze has cleared, and in the apartments across the courtyard I can see the small shapes of other people in their kitchens at this hour. The model is out. The work continues tomorrow.
I wash the glass. I turn off the kitchen light. I go to bed.