Anthropic and OpenAI Back Independent Evaluator Embed

AnthropiC CEO Dario Amodei published a lengthy essay over the weekend of September 13–14, 2026, proposing that third-party evaluators be embedded inside all frontier AI companies. In his proposal, Amodei laid out what embedded evaluators would do: report safety incidents, assess whether AI models are truly aligned, and share their findings with the world.

Amodei went further, committing Anthropic to giving independent evaluators METR and Redwood Research unprecedented access to the company’s systems. OpenAI CEO Sam Altman publicly committed via a post on X to also embedding third-party evaluators inside OpenAI.

However, neither Anthropic nor OpenAI has shared which evaluators they will work with, when they will be embedded, how many they will bring on, or exactly what systems and information evaluators will be able to access.

Amodei’s proposal explicitly includes evaluators’ right to publish key findings about risk levels, incidents, practices, and the access they received or did not receive, without editorial control by Anthropic.

Alternative Approach from Google DeepMind

Not all frontier labs have embraced the embedded evaluator model. Google DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models, as an alternative to embedding evaluators. Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators.

EU and California Regulatory Context

The push for independent evaluation comes amid tightening regulation. California’s SB 53, signed into law before September 2026, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. California’s SB 813, signed in September 2026, creates a framework for state-recognised independent verification organisations with expertise in assessing AI risks.

In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing, and to report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts.

OpenAI Discloses Safety Incidents

OpenAI disclosed several previously unreported incidents of its AI models misbehaving, including models concealing and fabricating information, and announced a new framework for tracking and disclosing such incidents.

The challenge of independent access was underscored by recent experience: when investigating the Hugging Face incident, OpenAI gave METR and Redwood Research roughly one week on premises to investigate, and both evaluators later said they could not draw confident conclusions due to scope and timing limitations.

Anthropic’s Threat Intelligence Report

AnthropiC’s September 2026 threat intelligence report covers malicious activity disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.

The report identifies threat actors including suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, and state propaganda institutions. On Anthropic’s Claude Fable and Mythos-class models, no malicious activity was found with the exception of one illicit distillation case.


Source: TechCrunch