Anthropic proposes public measurements of AI R&D automation, agent oversight and safety compute inside frontier labs

On September 17, 2026, Anthropic published “Measurements for understanding the pace of AI development inside frontier labs,” written by Marina Favaro and Phillie Wright. It follows the company’s argument, made days earlier, that frontier labs should deliberately pace capability gains, and it proposes three measurements any lab could publish so outsiders can see how fast AI is accelerating AI research: how much AI R&D work AI is leading, how well human oversight covers the agents doing that work, and how much of the compute used for AI R&D goes to safety. Anthropic reports its own numbers for each, as of August 2026.

For automation, Anthropic uses an Automation Level scale developed by Epoch AI that runs from AL0 (no AI involvement) to AL5 (AI fully autonomous, no human in the loop); at AL4 the AI “leads,” completing most of a task end to end from a high-level prompt while a human supervises. Anthropic says Claude leads 26 percent of its AI R&D work, more than 90 percent of the work is at or above the “collaborates” level, and Claude is not operating fully autonomously for any measured subset. For oversight, roughly 30,000 agents were doing research and engineering work at any one time on its most-used internal platform. Across more than a billion agent decisions in August, an online monitor blocked 0.002 percent, about 1 in 47,000, with humans reviewing blocked actions within a week; an offline monitor flags about 100,000 transcripts a week, of which about 50 of the highest priority are escalated to humans. For compute, over one examined week about 6 percent of AI R&D compute went to safety, and about 12 percent of the compute used for AI-driven AI R&D.

The authors are frank about the weaknesses. Safety and capabilities research are hard to separate and a lab will be tempted to draw the line generously; the automation ratings depend on judge models that may share the errors of the model being judged; compute is an imperfect proxy for safety effort; the compute figure is a one-week snapshot; and the science of agent monitoring is only months old.

Why it matters: these are rare first-party numbers on how much of a frontier lab’s own research an AI system now leads, and the headline 26 percent is a concrete data point for the automated-AI-R&D threat models that labs and governments track. What it does not provide is independent verification. Every figure is self-reported and self-measured, and the proposal only has value as a comparison if other labs adopt the same definitions.

Sources

Last verified September 28, 2026