Safety researcher Jacob Coxon resigned from Anthropic, publicly accusing the company and rival OpenAI of recklessly racing toward self-improving superhuman systems without adequate safeguards. Following his departure, Evan Hubinger, an AI safety lead at Anthropic, echoed Coxon’s concerns on social media. Hubinger stated he personally estimates a greater than 10 percent chance that AI could cause human extinction within the next decade, adding that the industry lacks a clear plan to ensure advanced alignment.

The public dispute underscores intensifying anxieties within frontier labs as development accelerates toward recursive self-improvement and autonomous code generation. The controversy emerges as major AI startups prepare for anticipated initial public offerings amid growing regulatory scrutiny over rogue autonomous agent incidents and model monitorability.

Why it matters

  • Signals growing internal talent strain and safety dissent within top frontier AI labs like Anthropic and OpenAI.

  • Highlights strategic alignment and governance risks for enterprise operators deploying increasingly autonomous, self-improving AI agents.

  • Could accelerate regulatory pressure and safety mandate enforcement on leading AI companies heading toward IPOs.

Source: theverge.com