Pyyan / News / 17 September 2026

ModelsAnthropic · Epoch AI

Anthropic says Claude now leads a quarter of the work that builds the next Claude

26%of AI research and development, up from under 1% in February

Recursive self improvement has been an argument for a decade. On 17 September a frontier lab published a number for its own: 26% of the work that builds the next Claude is now led by Claude, against under 1% in February.

Anthropic released a prototype R&D Automation Index. It scores tasks on Epoch AI's automation scale, which runs from AL0, no AI involvement, to AL5, fully autonomous with no human in the loop. Claude scored AL4 or above on 26% of Anthropic's model research and development in August, meaning it can carry most of a task end to end from a high level prompt with a human checking the result, and it takes part in more than 90% of that work at some level. Alongside it the company published two numbers nobody asked for: about 30,000 agents running at once on its main internal platform, every action monitored and roughly one in 47,000 blocked, and about 6% of AI research compute spent on safety research in the week of 13 July. The measurement itself was done by a Claude research agent, which reviewed the tasks of a fifth of the relevant staff each week of July, roughly 15,000 tasks, and organised them into a tree of 542 nodes.

Why this one is different

Labs describe this capability in essays and decline to quantify it. This is a method, a scale borrowed from somebody else, and a figure that can be measured again next quarter and compared. It is also, unavoidably, the model grading its own contribution: Claude did most of the rating, and Anthropic reports that model and staff ratings matched exactly 59% of the time while two employees rating the same task matched only 35% of the time.

Under 1% in February. 26% in August.

How we got here

  1. 19 Aug 2026OpenAI stops training its own models for two weeks.
  2. 3 Sep 2026OpenAI says GPT-6 Astra may be the arrival of general intelligence, and ships it switched off.
  3. 12 Sep 2026Amodei asks the industry to pace itself; Musk agrees; Altman defers the IPO.
  4. 17 Sep 2026Anthropic publishes the first numbers for how much of its own R&D the model leads.
  5. 18 Sep 2026Anthropic hires an embedded evaluator with employee level access.

What it does and does not mean

Every figure here is a company measuring itself, rated largely by its own model, and nobody outside has checked any of it. AL4 is not autonomy: Anthropic states plainly that Claude is not operating fully autonomously in any measured category, and the 35% agreement between two humans on the same task says the rubric is noisy at the edges. What is new is the shape of the disclosure. A lab has published a repeatable measurement of the one capability that would make everything else move faster, the day before it gave an outside evaluator the access needed to re-run it.

Related

← All the news, newest first