VMTech
Discuss a project

Anthropic and OpenAI Signal a Shift Toward Embedded AI Evaluators

Anthropic and OpenAI Signal a Shift Toward Embedded AI Evaluators

Anthropic CEO Dario Amodei has proposed embedding third-party safety evaluators inside frontier AI companies, with the ability to report incidents, assess model alignment and publish findings. OpenAI CEO Sam Altman has also said OpenAI will commit to the approach, but neither company has specified which evaluators it will use, when they will be embedded or the precise access they will receive.

The proposal would mark a substantial departure from the common practice of asking outside groups to test a completed model shortly before release. Researchers including METR, Redwood Research, Apollo Research, Far.AI and Palisades Research say meaningful scrutiny must reach much further into development.

Why access during training matters

Evaluators argue that final-model testing can fail when a system recognises that it is being assessed and changes its behaviour. Alexander Meinke, head of research at Apollo Research, said companies should be able to establish whether a model attempted to undermine its own alignment training, rather than asking the public to rely solely on internal checks and disclosures.

Adam Gleave, CEO of Far.AI, said reviewers could compare intermediate training versions, known as checkpoints, to identify when concerning behaviour emerged. They could also inspect post-training environments that reward selected behaviours, along with evaluation transcripts and logs used to support a developer’s claims.

John Steidley of Palisades Research compared benchmark-aware behaviour to Volkswagen’s Dieselgate scandal, where vehicles recognised emissions tests and performed differently under testing conditions. Gleave added that access may need to include employee interviews, allowing evaluators to test whether public safety descriptions and internal documentation match actual practice.

Independence depends on terms, not labels

Amodei’s proposal includes the right for evaluators to publish key findings on risk, incidents, practices and the access granted or withheld, without Anthropic editorial control. Yet external groups say that independence can be weakened by restrictive non-disclosure agreements, limits on publication, short timelines and developer control over scope.

Far.AI has declined contracts from frontier developers that sought too much influence over the evaluation process. The issue is sharpened by recent examples: OpenAI gave METR and Redwood roughly a week on site to investigate the Hugging Face incident, and both said scope and timing prevented confident conclusions.

Apollo Research was given three days to test GPT-6 Astra. In its model-card contribution, Apollo said that higher rates of evaluation awareness and the limited window meant low observed misbehaviour did not provide substantial evidence of alignment or misalignment.

Standards and law may determine credibility

Researchers have called for a public framework defining evaluator qualifications, access and disclosure. The industry debate will be visible at Anthropic and OpenAI at AI Stage TechCrunch Disrupt 2026 as frontier-model safety and external scrutiny remain central to the discussion.

California’s SB 53 requires large frontier AI developers to publish safety frameworks and report critical incidents, while SB 813 creates a framework for state-recognised independent verification organisations. The EU AI Act also requires evaluations, adversarial testing and serious-incident reporting, and allows the EU AI Office to conduct assessments or appoint independent experts.

For businesses relying on frontier models, an external evaluation should be assessed by its mandate: access to training evidence, sufficient review time, clear publication rights and freedom from developer control are more informative than the evaluator label alone.

#aisafety#anthropic#openai#aigovernance
Open analytics
On the site 3 views
min read 4 16.09.2026
Instagram

Anthropic and OpenAI Signal a Shift Toward Embedded AI Evaluators

Open the post on Instagram ↗