Anthropic’s Amodei proposes continuous evaluator access for AI firms

6 hours ago 73

Dario Amodei, CEO of Anthropic, has laid out a plan to embed independent third-party evaluators directly inside AI laboratories, giving them the kind of access typically reserved for full-time employees. The proposal, outlined in a September 12, 2026 essay titled “We Must Pace the Frontier,” represents one of the most concrete self-governance commitments an AI company has made to date.

What Anthropic is actually offering

Under the proposed model, external reviewers would receive desks at Anthropic’s offices, company access badges, and laptops. They would function with a level of integration typically associated with internal staff, maintaining ongoing access to the company’s AI development environments rather than parachuting in for periodic check-ups.

Perhaps the most notable piece of the commitment: evaluators would be able to publish their findings with minimal redaction and without Anthropic exercising editorial control over the results.

Amodei specifically named METR, an AI safety evaluation organization, as a potential embedded evaluator under this framework.

The shift from episodic to continuous evaluation is the key architectural change here. Traditional third-party audits capture a snapshot of a system at a particular moment in time. Embedded evaluators would instead observe the development process as it unfolds, catching potential issues before they get baked into deployed models rather than flagging them after the fact.

Why this matters beyond Anthropic’s walls

Anthropic has long positioned itself as the safety-first AI lab, a reputation it has cultivated since its founding by former OpenAI researchers. The company has previously advocated for mandatory independent evaluations and pre-deployment testing of high-risk AI capabilities. This latest proposal takes those positions and converts them from talking points into operational commitments.

Amodei’s essay frames the embedded evaluator model as a way to create “verifiable safety commitments” in an environment where trust in self-reporting alone is increasingly insufficient.

The competitive and regulatory calculus

There are legitimate questions about how this works in practice. Access badges and laptops are concrete, but the boundaries of what evaluators can actually examine, how deeply they can probe proprietary training methods, and whether “minimal redaction” stays minimal when commercially sensitive information is at stake will determine whether this is a genuine governance innovation or an elaborate PR exercise.

The AI safety community will likely watch the implementation closely, particularly the degree to which METR or other evaluators can operate without implicit pressure. The contract terms governing publication rights will be scrutinized heavily, because the gap between “can publish with minimal redaction” and “can publish anything meaningful” can be wide.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article