Claude Now Leads 26% of Anthropic's R&D: Inside the Index That Measures How Much AI Research Is Done by AI
Anthropic has disclosed for the first time how much of its own AI research is done by its own model. Its R&D automation index puts the figure at 26%: Claude "leads" that work, more than 90% of it involves the model in some form, and no measured task reaches full autonomy.
What Anthropic Published and What "Lead" Actually Means
The piece is titled "Measurements for understanding the pace of AI development" and it is published by the Anthropic Institute. In it the company introduces the Anthropic R&D Automation Index, a prototype for measuring what share of its model research and development work is performed by its own models. It does not measure capabilities: it measures the process that produces them.
A Prototype Index, Not a Results Report
The construction is explicit, and it is worth understanding before trusting the number. Anthropic catalogued every kind of AI R&D task performed inside the company, rated how automated each one currently is, and weighted the result by how much staff time it consumes. The sample came from 20% of the staff in the departments that make up the development loop, week by week through July 2026: roughly 15,000 granular tasks, later organised into a tree of 542 nodes, 378 of which are concrete leaves such as diagnosing defects in an evaluation platform. That tree is frozen so every later measurement runs against the same basket of work.
The Difference Between Leading and Collaborating on a Task
"Leads" is a level on the scale, not a headline metaphor. Collaborating means handling large chunks of a task under close human direction. Leading means completing most of the work end to end from a high-level prompt, with a person supervising and deciding whether the output ships. The document uses an infrastructure example: a nightly pipeline that breaks. At the collaborating level the person stays in the conversation, answers when the model hits something new and reviews the change line by line. At the leading level the person hands over the failure alert; the model reads the logs, finds the cause, writes and tests the fix, handles surprises and writes up the report — then the person reads it, asks whatever they want and decides whether the change ships tonight or waits.
Read also
The AL0–AL5 Scale, Explained Without Jargon
The six levels come from a framework developed by Epoch AI and read like a ladder of delegation:
- AL0: no AI involvement.
- AL1: minimal AI involvement.
- AL2: AI assists.
- AL3: AI collaborates, under close human direction.
- AL4: AI leads, with human supervision.
- AL5: AI operates fully autonomously, with no human in the loop.
Where Claude Stands Today
In the August 2026 measurement, Claude "leads" 26% of Anthropic's AI R&D work, and more than 90% of that work sits at "collaborates" or above. The company is explicit about the other end of the scale: it is not operating fully autonomously for any measured subset. AL5 appears as the direction the scale points in, not as a result achieved.
The Numbers: From Under 1% to a Quarter of the Work
What makes this news is not the 26% on its own, but the progression behind it:
- February 2026: under 1% of R&D work.
- May 2026: 12%.
- July 2026: 22%.
- August 2026: 26%.
In seven months, AI moved from the margins of the process to covering a quarter of it. The "collaborates or above" bucket already covers more than 90%, which describes a company where nearly all development work passes through a model at some point.
Inside the Agents: 30,000 at Once and a Monitor That Rarely Steps In
The second set of measurements does not look at the model, it looks at the agent scaffolding. As of August 2026, Anthropic's most-used internal platform runs roughly 30,000 agents doing research and engineering work at any one time. Every one of their actions passes through an online monitor before execution and is ingested by an offline monitor afterwards.
How Often the System Actually Intervenes
Across more than a billion decisions analysed during August, the monitor blocked 0.002% of them, about 1 in 47,000. The review system flags one to two transcripts in every thousand for follow-up, on the order of 100,000 a week, and only about 50 per week reach human review. Anthropic presents those figures as a sign that oversight is not currently saturated. A critical reader will also see the underlying question there: if the block almost never fires, how much oversight is genuinely happening and how much is installed capacity.
The Fine Print: A Self-Reported Number
All of the above is measured and published by Anthropic about itself. No independent verification has been published, so it should be read as a lab's number about its own work, not as an audited figure.
Who Measures and Who Has Not Verified Yet
The index relies on classifications made by the company's own models, and the document acknowledges the obvious risk: the model acting as judge could make the same kinds of errors as the model being judged. To calibrate it, Anthropic asked internal owners of each work area to rate the automation level of their areas without seeing the evidence the models had gathered. Exact agreement between model and human was 59%, and the two ratings sat within one level of each other 97% of the time. For reference, agreement between two humans was 35%. The text itself concedes that borderline cases — where "collaborates" ends and "leads" begins — remain genuinely arguable.
What an Independent Evaluator Would Ask
Three concrete things: what counts as a task and how it is weighted, who checks each task's classification on the scale, and whether the system that detects unwanted behaviour is independent of the one measuring progress. The company says it plans to embed external evaluators from multiple organisations, with access to internal processes, systems and data comparable to what its own risk teams get. Until that happens and is published, the 26% is a figure without outside audit.
What It Means If You Build With AI
There is a detail that tends to get lost and that does transfer outside Anthropic: the company publishes a third measurement, compute. During the week analysed, about 6% of the compute going to AI R&D and roughly 12% of the compute going to AI-driven AI R&D was allocated to safety work. Those are conservative, single-week estimates, but they raise the question any team should ask before delegating: of everything you hand to an agent, how much budget and time do you spend reviewing what it does.
The metric also teaches you to read an agent's work without exaggerating it. "Leads 26%" sounds like full autonomy, and it actually describes effective supervision over a quarter of the work. Human oversight is not decoration on the system: it is the component holding up the other 74%.
Conclusion
Anthropic published an awkward measurement and published it with its methodology in view, which is more than labs usually do: a quarter of its R&D already runs with the model in front and humans supervising, after going from almost nothing to 26% in seven months. The open question is not whether the number is high, but who audits it. The wider debate already has context on the blog: there are the six misalignment cases OpenAI published, Dario Amodei's call to slow the pace, and the MCP standard, which explains the layer where much of this delegation happens.


