Yesterday, Anthropic published three measurements of how fast AI is being built inside its own lab: how much of its AI R&D is performed by AI, how well its AI agents are overseen, and how its compute is allocated. As of August, Claude leads 26% of Anthropic’s AI R&D work, which is up from under 1% in February.
Anthropic scores the work on a six-level scale from Epoch AI that runs from AL0 (no AI involvement) to AL5 (fully autonomous, no human in the loop). At AL4 (AI leads), Claude completes most of a task end-to-end from a high-level prompt while a human supervises. More than 90% of the company’s AI R&D work is at AL3 (AI collaborates) or higher. None of it is at AL5.
In August, about 30,000 agents were doing research and engineering work at any one time on Anthropic’s most-used internal platform. Every action passes through an automated monitor before it executes. Across more than a billion decisions, that monitor blocked 0.002% (about 1 in 47,000). A second monitor reads everything after the fact and flags roughly 100,000 transcripts a week; about 50 of those reach a human reviewer. The company says that during one week in July, about 6% of the compute Anthropic spent on AI R&D went to safety work.
A Claude agent built the list of roughly 15,000 R&D tasks from Slack and internal documents, and a separate Claude judge assigned the automation ratings. Claude graded Claude. Anthropic acknowledges the judge “could make the same kinds of errors as the model it is checking,” and says it will embed independent evaluators from multiple organizations to verify the numbers.
Recursive self-improvement is a fancy word for a loop, and this is the first time a frontier lab has told us how much of the loop is closed. The post opens with a link to Dario Amodei’s call to pace the frontier. The frontier labs don’t need anyone’s permission to slow down. (See: Weapons Grade AI and the Licensed Frontier). I’m a bit skeptical of a company publishing the yardstick by which it wants to be measured. It’s how regulatory capture starts.
One important best practice to note: Anthropic gives every agent its own identity, ties every record to that identity, and makes agents communicate in the open where monitors can read along. We advise our clients to do the same thing. Agents are synthetic employees, and they need badges. Coverage, review latency, and escalation rate are three numbers you should be able to report for every agent you run.
Every company needs a Claw strategy. Do you have one?
Author’s note: This is not a sponsored post. I am the author of this article and it expresses my own opinions. I am not, nor is my company, receiving compensation for it. This work was created with the assistance of various generative AI models.
About Shelly Palmer
Shelly Palmer is the Professor of Advanced Media in Residence at Syracuse University’s S.I. Newhouse School of Public Communications and CEO of The Palmer Group, a consulting practice that helps Fortune 500 companies with AI strategy, implementation and governance, as well as technology, media and marketing. Named one of LinkedIn’s Top Voices in Technology, he is a bestselling author, covers tech and business for Fox 5’s Good Day New York, is a regular commentator on CNN, and writes the popular daily business blog Think About This. Follow @shellypalmer or visit shellypalmer.com.