Who checks the code before it ships has changed twice in the past two decades and is changing again. Microsoft’s SDET era put quality in a specialist role. Continuous delivery moved much of that responsibility into the product team and its automated systems. Coding agents now write, test, and review a growing share of each change.
That history points to a more useful question than “Can AI review code?” The question is: what evidence must a change produce before it earns the right to reach more users?
The answer in Gergely Orosz’s reporting is remarkably consistent. Fast teams still demand evidence. They encode more of it in tests, rollout systems, telemetry, and rollback paths. Agents can run those mechanisms and help create them, but they do not remove the need for an independent signal—or for a person and team accountable for the result.
1When testing was a department
How Microsoft does Quality Assurance traces the Software Development Engineer in Test role that Microsoft pioneered in the 1990s. SDETs wrote test code and built testing infrastructure alongside developers. A 2:1 ratio of developers to SDETs was common until roughly 2014.
The arrangement produced expertise and clear investment in testing. It also produced queues. A developer finished a feature, handed it to test, waited for feedback, fixed the defects, and handed it back. Orosz calls this “ticket ping-pong.” Quality had an owner, but that owner often entered the loop after implementation.
The split was visible in everyday design decisions. On Orosz’s Skype for Xbox One team, six SDETs served twelve developers. The test group owned manual verification, integration and end-to-end automation, milestone test plans, and tricky performance benchmarks. Even unit tests became a boundary dispute: some developers saw all automation as test-team work; others argued that tests this tightly coupled to the code belonged with its author. A specialist team made outsourcing quality feel locally rational.
On Skype for Web in 2014, a team shipping daily quietly removed that boundary. SDETs began taking production work; developers became responsible for their own tests. Microsoft retired the SDET title company-wide later that year. The organizational change mattered, but the technical rewrite mattered more.
Visual Studio Team Services supplied the hard version of the story. In 2015, its tens of thousands of tests took most of a day to run, more hours to classify false failures, and sometimes days or weeks to repair after a legitimate product change. The combined engineering organization discarded much of that eight-year-old suite and rebuilt around smaller tests. Within two years, the mix had inverted: granular unit and integration checks replaced end-to-end tests as the foundation.
This history is easy to overgeneralize. Dedicated QA did not disappear from software. How Big Tech does QA found that Apple retained substantial QA investment and Amazon kept an SDET career path, especially around hardware and slower-release products. Google’s Engineering Productivity organization still supplies testing expertise and common infrastructure even where product teams own outcomes.
Nor does deleting a title magically transfer the capability. Quality Assurance Across the Tech Industry, based on responses from 47 tech professionals, records what happened after Indeed eliminated most QA roles in its 2023 layoffs: one engineer reported inheriting test systems whose builders were gone and a sharp fall in test quality. Other respondents described small startups that consciously accept minor outages but invest in recovery and protecting user data. The viable model depends on product risk—but the responsibility cannot simply vanish with the org chart.
The real handoff was narrower and more important: on fast-moving software teams, quality stopped being a phase at the end of development and became a property of the development system.
2When quality became a system
Once developers owned quality, automation had to make that ownership sustainable. Inside Stripe’s Engineering Culture, Part 1 describes the stack: automated tests, CI/CD, monitoring, code review, progressive rollout, and blue-green deployment. Its second part supplies the scale. Stripe’s more than 50 million lines of code could be checked in about 15 minutes by tests that would take 50 days on one CPU. In 2022, its core payments APIs were deployed 5,978 times; 1,100 deployments failed acceptance criteria and were rolled back automatically.
That is not a story about tests making failures impossible. It is a story about a system making failed changes cheap to detect and reverse.
Stripe turned that system into a product for its engineers. Its deployment console puts the active deployment, queued changes, pending pull requests, live traffic, dashboards, logs, and runbooks in one view. Critical services keep the old and new versions live through blue-green deployment, moving traffic only after checks pass. User-facing changes require two-person signoff; feature rollouts can begin with just 0.0015% of traffic before advancing through explicit stages.
Shipping to Production generalizes the pattern: code review and automated tests before release; canaries, feature flags, staged rollouts, and controlled exposure during release; monitoring and rollback after release. Meta’s four-stage canary process once caught an ads defect at stage two, before broad exposure. OpenAI’s earlier approach to ChatGPT, described in How does ChatGPT Ship So Quickly?, similarly favored monitored, incremental releases over a big-bang launch.
Meta’s sequence shows what “release” now means. A change first passes automated checks, then reaches employee dogfooding, then a small regional market, and only then the full user base. One ads defect was caught in the second stage, while employees—not customers at internet scale—were the blast radius. OpenAI used the same logic for early ChatGPT releases: each monitored increment was both product learning and a safety boundary.
The operating model also changes accountability. Amazon’s engineering culture is explicit: “you build it, you own it”. The engineer on call is not merely paged; the surrounding escalation path makes service health a management concern. Healthy Oncall Practices explains why the feedback loop matters: the team with the most context responds fastest, and the people making design decisions feel their operational consequences. This only works if on-call is staffed, humane, and given time to improve the system; otherwise ownership turns into burnout.
The lesson is not to copy Stripe, Meta, Amazon, OpenAI, or Google wholesale. Adopting Software Engineering Practices Across the Team warns that every practice carries costs and must solve a defined problem. Google’s common tooling and processes, described in Inside Google’s Engineering Culture, work at Google’s scale and in Google’s context. A release gate is useful only when it produces information the team can act on.
Three properties make a check a real gate:
- It is independent. The check does not merely repeat the author’s assumptions.
- It controls exposure. A failed signal can stop or reverse the rollout.
- It has an owner. Someone is accountable for interpreting the evidence and improving the system.
3Agents raise output faster than confidence
AI changes the economics of producing a change. It does not automatically change the evidence required to trust one.
How Claude Code is built reported that about 90% of Claude Code’s code was written with Claude Code, while the team averaged roughly five releases per engineer per day. Anthropic also reported a 67% increase in pull-request throughput while its engineering team doubled. Slow down to speed up adds broader vendor data: among Cursor users, code output rose 2.5 times and pull-request size tripled over 18 months, while more changes were accepted without human review. These figures describe activity in specific populations; they are not, by themselves, measures of reliability or business value.
The most revealing Claude Code story is not a benchmark. For a todo-list feature, product lead Boris Cherny generated roughly twenty prototypes in a few hours across two days. After each pass he tried the interface, adjusted it, shared promising versions with colleagues, and threw away the ones that felt wrong. The agent made implementation cheap enough that product judgment—not typing—became the scarce step.
The same team automated the small work without pretending it had eliminated review. A new GitHub issue receives an agent’s first triage and often an attempted fix; Orosz reports that the first shot works about 20–40% of the time. Claude writes most of the project’s tests and performs the first code-review and security-review passes. At the time of the reporting, a fellow engineer still performed the second review. Automation increased the number of checks, but it did not collapse maker and checker into the same act.
The full account in How building software is changing at Anthropic makes the distinction concrete. Bun creator Jarred Sumner used 64 parallel agents and, at API prices, about $165,000 in tokens to port more than 535,000 lines of Zig to Rust in 11 days. Only about 15% of the time went to initial implementation; the other 85% went to compiling, fixing, testing, and verification.
The headline speed depended on conditions that are easy to omit:
- Sumner was the codebase’s deepest domain expert.
- Bun already had a robust, language-independent test suite.
- The migration started with a detailed plan and style guide.
- Separate agents performed review, security scanning, and fuzzing.
- Pull requests still passed explicit quality gates before a person merged them.
Sumner did not try to read every generated line. He built evidence around the change: eleven security-scanner runs, agent-written fuzzers, tests executed outside the coding session, and explicit merge gates. That is the AI-era release gate in miniature. The agent generates more change, but the system demands more machine-checkable evidence in return.
The full Anthropic reporting also resolves an apparent contradiction about planning. Some product teams replace PRDs with rapid prototypes. A six-month platform project such as Claude Managed Agents still used a PRD because it had to align many teams and three cloud providers. The dividing line is not “before AI” versus “after AI.” It is coordination cost, maturity, and blast radius.
Impressions from visiting OpenAI, Anthropic, & Cursor pushes the idea further: a growing share of engineering work is designing the environment in which agents operate. Cloud agents need isolated machines, durable state, tools, context, and ways to surface trouble when no person is watching. Cursor discovered that an asynchronous agent has no natural moment to complain to a human, so the team periodically asks agents to “confess” problems from their runs and routes those reports to the infrastructure team. Even the feedback loop must now be engineered.
The warning comes from AI’s impact on software engineers in 2026, Part 2, which summarizes more than 900 survey responses. Respondents reported both faster iteration and declining code quality. The cleanest formulation was that AI amplifies the engineering culture already present. Strong guardrails compound; weak ones are bypassed faster.
4What slips through
More of Orosz’s 2026 reporting describes the same imbalance from the other side: code production accelerates first, while review attention, product testing, and operational guardrails struggle to keep up.
The June 2026 Meta incident is a caution, but it needs careful attribution. In Slow down to speed up, Orosz reports—based on conversations with Meta engineers—that an AI-generated, AI-reviewed change allowed Meta AI to change the email address on another person’s account. He also reports that relevant trust-and-safety staffing and on-call coverage had been weakened.
That account does not prove that AI review alone caused the incident. It describes a correlated failure of generation, independent review, security oversight, and operational ownership. Treating it as merely “the model wrote a bug” would miss the more useful lesson: several layers that should not share the same failure mode did so at once.
In March, Are AI agents actually slowing us down? opened with a smaller but unusually legible failure. On Claude.ai, a visitor could begin typing before subscription data finished loading; when that data arrived, the input reset and erased the first words. The defect sat on the flagship landing page, reproduced on every visit in the reported flow, and affected paying customers until a public complaint reached the team. Nothing exotic failed. A basic end-to-end user journey simply had no effective owner or alarm.
The same article reported a higher-blast-radius case from late 2025: AWS’s customer cost calculator was interrupted for thirteen hours after engineers allowed the Kiro coding agent to make changes and it chose to delete and recreate the environment. Amazon’s retail organization, meanwhile, responded to a rise in AI-associated incidents by requiring senior approval for AI-assisted changes from junior and mid-level engineers. The important control was not “use less AI”; it was restoring authority in proportion to risk.
By July, The Pulse: New trend—concern about massive increase in code review load reported that engineering leaders were treating review as the new bottleneck. Uber built smart assignments and change-risk profiles into Code Inbox; Cloudflare, Faire, and HubSpot built their own review tools. Yet the human failure mode remained: some engineers approved when the AI reviewer had no comments, while colleagues who continued reviewing with full attention were buried in oversized, machine-produced pull requests. A second model can add a signal. It cannot manufacture attention.
The 2023 Datadog outage shows the same systems principle without an AI component. Inside Datadog’s $5M Outage describes a security update that automatically reached tens of thousands of virtual machines in a narrow time window. Restarting systemd-networkd removed network routes, including routes needed by the control plane that could have repaired the damage. The first internal alert fired three minutes after impact; a customer-visible incident was declared at 31 minutes; full recovery, including backfill, took nearly 48 hours. Datadog later disclosed about $5 million in lost revenue.
Crucially, neither of the two security fixes in that update directly caused the outage. The failure came from the update path: an overlooked legacy auto-update channel, synchronized rollout, and a circular dependency that widened the blast radius.
This is why Incident Review and Postmortem Best Practices argues for systemic investigation instead of stopping at the person or line of code nearest the failure. The article’s survey of more than 60 teams found broadly similar incident processes, but its best examples focus on reconstructing how people understood the system at the time. Honeycomb’s reviews ask who knew what, when, and how—not only which action items can be extracted.
Migrations Done Well makes the preventive version of the same point. A migration needs dedicated monitoring, shadow traffic where practical, validation against production-shaped data, dry runs, and an explicit rollback strategy. The hardest changes need stronger evidence because their test oracle is weaker and their rollback path is often narrower.
5The modern release gate
The articles point to a practical design for teams shipping agent-generated code:
- Require evidence with the change. Tests should fail without the fix and pass with it. Security-sensitive changes need security-specific checks. Generated tests should be reviewed against requirements, not trusted because they are green.
- Separate maker and checker. Use a fresh context, a different model or tool, deterministic analysis, and human review where risk warrants it. Independence matters more than whether the checker is human or AI.
- Classify risk before choosing the path. Low-blast-radius changes can earn automated merging. Authentication, payments, permissions, data migrations, and infrastructure control planes deserve stronger review and slower exposure.
- Expose progressively. Use test environments, shadow traffic, canaries, feature flags, or staged rollouts so production becomes a source of bounded evidence rather than an all-or-nothing bet.
- Make failure cheap to reverse. A gate without a practiced rollback path only detects regret. Measure time to detect and time to mitigate, not just whether CI passed.
- Keep a named human owner. Agents can execute checks and recommend a decision. The team still defines acceptable risk, owns on-call, and repairs the system after a miss.
The fastest teams in these reports do not ship because they trust agents more. They ship because they have made trust operational: a change arrives with evidence, reaches users in stages, is watched by independent signals, and can be reversed quickly.
The gate is moving. Accountability is not.