Anthropic published operational metrics tracking its internal use of AI in building future models, reporting that Claude “leads” 26 percent of its research work as of August 2026. Using a scale developed with Epoch AI ranging from AL0 (no AI) to AL5 (full autonomy), Anthropic classified tasks where Claude independently resolves bug reports or code fixes under human oversight as AL4 (“AI leads”), up from under one percent in February.

The classification process relied heavily on self-evaluation, with Claude models analyzing internal Slack logs and documentation to assign levels. Anthropic acknowledged ambiguity in the scoring system, noting that human employees agreed on automation levels for the same task only 33 percent of the time, while Claude matched human ratings 59 percent of the time.

Additionally, the 26 percent figure measures human work hours associated with tasks rather than critical strategic decision-making. Anthropic also revealed that its internal agent platform runs approximately 30,000 concurrent agents overseen by automated real-time monitors, which intervened in 0.002 percent of over one billion actions in August.

Why it matters

  • Demonstrates rapid adoption of AI agents in core software development, with over 90% of tasks assisted by internal LLMs.

  • Exposes methodological challenges in evaluating model autonomy and the limitations of LLM self-grading systems.

  • Signals growing operational reliance on real-time agent monitoring systems for high-volume enterprise agent deployments.

Source: the-decoder.com