Anthropic sets out independent oversight plan for frontier AI

Anthropic commits to embedded external evaluators
Anthropic CEO Dario Amodei has set out a three-part approach to “pace the frontier” of artificial intelligence development, starting with a unilateral commitment to embed independent evaluators inside the company. The proposal follows a period of sharply accelerating AI capabilities and the OpenAI-Hugging Face hack, which Amodei cited as reasons for a more deliberate approach.
Amodei wrote that AI progress will still appear rapid, but argued that companies must slow the rate at which they improve model capabilities and use the time gained wisely. His first proposal is for third-party organisations, including METR, to place evaluators within frontier AI companies to verify compliance with pacing and safety commitments and to ensure safety incidents are reported.
Anthropic says those evaluators would receive company badges, desks and laptops. Their access would be mostly comparable to that of internal risk-assessment teams, subject to exceptions required by law or contract. Amodei compared the arrangement with regulators embedded alongside bank employees and called on governments to require similar access at other frontier AI firms.
Government mediation and international constraints
The second strand calls for leading AI companies in democratic countries to coordinate common safety standards and limits on unchecked progress. Amodei said US government mediation or enablement would be useful because antitrust rules may otherwise constrain narrowly focused safety discussions. The debate also arrives as Anthropic and OpenAI at TechCrunch AI Stage puts Anthropic and OpenAI in the same public AI industry forum, while their executives have differed over the pace and governance of model development.
Amodei addressed a frequent objection to slowing US development: the prospect of Chinese AI leadership. He argued that refusing sales of powerful chips and semiconductor manufacturing equipment to Chinese companies, alongside action against model distillation, could slow China’s progress enough to widen America’s lead over the next three to five years.
His third proposal is global coordination involving the United States, allies and, where possible, authoritarian governments including China. He acknowledged strict limits on such cooperation, but suggested that narrow agreements could still cover obviously dangerous applications, such as producing biological weapons or allowing users to do so.
A debate shaped by trust and current risks
The post follows intensified argument over AI safety after researcher Jacob Coxon said he was leaving Anthropic, alleging that leading AI companies were taking extreme risks. Amodei did not explicitly address the resignation, but said public skepticism towards technology companies, the wider industry and government has created a crisis of trust.
Critics have questioned whether catastrophic-risk warnings distract from harms already associated with AI, and some have warned that frontier-model regulation could favour the largest providers. Amodei maintained that AI can greatly improve human life, while arguing that those benefits depend on building and deploying the technology in the right way. For businesses, the immediate implication is to make independent evaluation, documented incident reporting and accountable governance part of AI adoption decisions rather than treating them as optional safeguards.

