Anthropic disclosed Wednesday that its AI system Claude generated over 80% of the code merged into the company’s production codebase in May 2026, up from low single‑digit percentages when Claude Code debuted in February 2025. The shift has already transformed the workflow of Anthropic’s engineers. In the second quarter of 2026, a typical engineer merged eight times as much code per day as in 2024, and an internal poll of 130 research staff said the latest model, Mythos Preview, boosted their output roughly fourfold compared with work done without AI assistance.

Claude’s performance on the most demanding, open‑ended engineering tasks also improved dramatically. The model’s success rate climbed to 76% in May 2026, a 50‑percentage‑point increase over six months. Anthropic cited a recent incident where a routine upgrade caused tens of thousands of training jobs to crash. An engineer fed Claude the live incident details and cluster access; within two hours Claude isolated an obscure debugging flag, reproduced the crash, and confirmed a fix that would normally take two to three days.

Quality gaps are narrowing, too. Staff now rate Claude‑written code as roughly on par with human‑written code, up from “somewhat worse” in late 2025. An automated Claude reviewer scans every proposed change before it merges, and a retrospective analysis suggests it would have caught about one‑third of the bugs behind past Claude‑related incidents before they reached production.

Beyond coding, Anthropic is testing Claude’s research capabilities. In April 2026 the company released a demonstration in which nine parallel Claude agents tackled an open‑ended AI safety research problem. Over 800 cumulative hours and roughly $18,000 in compute, the agents recovered 97% of the performance gap on the task, while two human researchers achieved only 23% in a week. Another internal experiment measured Claude’s ability to choose the better “next step” during research. In November 2025 Claude matched human judgment 51% of the time; by April 2026 that figure rose to 64%.

The rapid progress aligns with broader trends tracked by METR, a nonprofit that benchmarks AI capabilities. According to METR, the length of tasks AI can reliably complete on its own has been doubling roughly every four months. Claude Opus 4.6 now handles 12‑hour tasks, and Mythos Preview can sustain work for at least 16 hours—far beyond the hour‑and‑a‑half tasks managed by Claude Sonnet 3.7 in early 2025. If the curve continues, tasks that currently require days of skilled human effort could become routine later this year, with week‑long tasks possibly arriving by 2027.

The downstream effects are already visible on platforms like GitHub. The code‑hosting service processed about one billion commits in 2025; by mid‑2026 it was handling 275 million commits per week, on track for 14 billion annually. Claude Code accounts for 4.5% of all public commits on GitHub, generating roughly 2.6 million weekly. As Claude produces more code, human code review has become the new bottleneck, a textbook illustration of Amdahl’s law.

In the accompanying Anthropic Institute paper, the company pivots from productivity gains to a call for a verifiable global mechanism to slow or temporarily pause frontier AI development. Anthropic argues that a unilateral pause by a single lab would merely shift leadership, whereas a coordinated, verifiable agreement among multiple labs and nations could buy time to address the technology’s “immense implications.” The paper draws parallels to nuclear arms‑control treaties but notes the unique challenges of concealing AI training runs and the massive incentives to defect.

Anthropic frames the issue as a three‑scenario outlook. The first scenario sees the current trajectory stall, still reshaping the economy. The second envisions AI‑driven automation of development while humans steer research direction, allowing tiny teams to match the output of massive organizations. The third, most speculative scenario, predicts full recursive self‑improvement, where AI systems design and train their own successors. While Anthropic admits it lacks solid intuition about the third outcome, it cautions that even a recursively intelligent system cannot accelerate every domain—drug effects, constitutional processes, or personal relationships remain bounded by external constraints.

The paper arrives as Anthropic expands its enterprise offerings, selling Claude as a productivity revolution while simultaneously warning that the same acceleration could demand an “emergency brake.” Whether the warning reflects principled transparency or strategic positioning will become clearer as the industry’s pace continues to outstrip existing oversight mechanisms.

This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.