The Rundown AI homepage
Artificial intelligence/News & analysis

Anthropic says Claude leads 26% of its measured AI research

Anthropic says Claude led 26% of its measured AI research in August. Its call to pace AI progress now faces a test of access and public transparency.

By The Rundown Editorial TeamReviewed by Kelly Pitts3 min read
Claude now leads 26% of Anthropic's AI research — newsletter story image
Image source: Anthropic

Anthropic has published internal estimates that put Claude at the “leads” level for 26% of its measured AI research and development work in August, up from under 1% in February. Humans supervise that work.

The figures, covered in The Rundown, give the public a closer look at how much AI now helps build AI. They arrive after CEO Dario Amodei called for pacing AI progress earlier this month.

How Anthropic measured Claude’s role

Anthropic applied Epoch AI’s scale to a fixed set of tasks, weighted by the human time they take. “Leads” means Claude completes most of a task under human supervision. It reached “collaborates or above,” which includes work under human direction, for over 90% of the measured work. Full autonomy was zero. That level would mean finding a problem, fixing it and shipping the result to production without human involvement.

Claude itself rated the tasks, and its scores matched human ratings exactly 59% of the time. The ratings leave room for disagreement about where work falls on the scale.

On its largest internal platform, Anthropic reported about 30,000 agents running at once. It said all their actions entered monitoring, with roughly one in 47,000 decisions blocked across more than one billion decisions in August. The monitoring claim covers that platform.

Safety work received about 6% of R&D computing resources during July 13–20. That snapshot excluded safeguards classifiers.

Why it matters

On September 12, Amodei warned that AI helping build its successors could speed development beyond people’s ability to understand and control the resulting systems. His immediate commitment was to bring outside evaluators into Anthropic with access similar to employees. Broader limits and shared standards would depend on coordination among companies and governments.

These measurements make that call for restraint more concrete. They expose how far automation has progressed inside the company asking others to slow the pace. For researchers and safety teams, human supervision still leaves a hard question about whether review can keep up with more experiments. Anthropic argues that research can accelerate while people continue choosing goals and models execute the work. It says fully autonomous development of a successor has yet to happen and is not inevitable.

Public transparency also depends on who can test the numbers. Amodei acknowledges that companies choose what their own disclosures include and omit. Evaluators would need to inspect underlying records, challenge classifications and report adverse findings for outsiders to judge whether the measures hold up.

The rate of blocked decisions leaves open how many problems went undetected. Anthropic’s August risk report, covering conditions as of July 15, said the company had yet to evaluate its automated offline monitoring from start to finish. Evidence about detection quality would help evaluators assess August’s platform monitoring.

Anthropic has taken a concrete step. On September 18, it announced an evaluation partnership with Accenture, led by its Faculty business, covering red teaming, alignment assessments and safeguards testing. Anthropic will fund the work directly, and the announcement left access and reporting rules unsettled. Those rules will help determine the evaluators’ independence and what their work reveals to the public.

Competitors have offered different levels of support. TheWrap reported that Altman endorsed pacing and said OpenAI would give independent evaluators access similar to employees. Musk’s statement was “Dario is right.” Whether those endorsements produce comparable access and public reporting across labs remains an open question.

Sources & further reading

This story builds on reporting from The Rundown newsletter on September 24, 2026.