Sunday Sep 20

Claude Leads 26% Of Anthropic's Own R&D

20SEP
1% TO 26%CLAUDESUPERVISOR

Anthropic measured how much of its research Claude runs on its own. Answer: 26% of AI R&D tasks, up from under 1% in February. Humans still supervise every one.

The number comes from Anthropic's new Institute. Jack Clark, co-founder, directed the work. They sampled 15,000 real staff tasks and had Claude grade each one on a five-level autonomy scale.

Level 4 means Claude does most of the task end to end and a human checks. That covered 26% in August. Over 90% of tasks now sit at collaborate or higher. Fully autonomous: zero categories.

The scale is the story. About 30,000 agents run at once. A monitor reviewed over 1 billion actions in August and blocked 1 in 47,000. Safety got about 6% of compute.

full brief & sources

⚡ Why this matters

  • First time a frontier lab published a measured number for how much of its own research AI does. Everyone else talks in vibes.
  • The trend line matters more than the number: under 1% in February to 26% in August. That is a six-month curve, not a decade.
  • It reframes the recursive self-improvement debate from thought experiment to a dashboard metric.

🔍 What happened

  • Anthropic Institute report by Marina Favaro and Phillie Wright, research direction from co-founder Jack Clark, published Thursday.
  • Method: sampled about 15,000 tasks from 20% of staff via Slack and docs, sorted into a 542-node task tree, graded by a Claude judge against Epoch AI's autonomy levels.
  • The Claude judge agreed exactly with humans 59% of the time. Humans agreed with each other only 35%. Within one level: 97%.
  • Ops numbers: about 30,000 agents at a time, over 1 billion monitored decisions in August, 0.002% blocked, about 50 transcripts a week reach human review.
  • Compute snapshot for one July week: about 6% of AI R&D compute went to safety work. Anthropic says third-party evaluators will be embedded next.

💬 Smart takes

  • Bloomberg framed it as Claude 'driving' a quarter of R&D. Anthropic's own wording is more careful: Claude leads, humans supervise.
  • The Neuron asked the sharp question: who sets the metric? A lab grading its own AI with its own AI is a conflict of interest by design.
  • Quartz and Engadget both flagged that level 5, fully autonomous, is zero across all categories. The headline number is about delegation, not replacement.

🧭 Where this goes

  1. LikelyOpenAI and Google DeepMind publish comparable autonomy numbers within two quarters. This becomes a new benchmark race.
  2. Possibleoutside evaluators get access to the task tree and the judge, and the 26% gets revised, in either direction.
  3. Wild Carda regulator asks for this metric as a disclosure requirement, the way emissions get reported.

🥄 The Spoon Take

This is the most useful safety document of the year, and it is not about safety. It is a productivity audit that happens to show how fast the loop is closing. If your team still debates whether AI can run research, Anthropic just handed you a chart. Watch the slope, not the 26%.

🤔 Pushback

A Claude judge grading Claude's autonomy, on tasks Anthropic chose, is not an independent measurement.