AI Evaluator Forum urges independent oversight for AI safety evaluations

1 hour ago 27

If you’re going to grade someone’s homework, you probably shouldn’t be sitting in their living room, eating their snacks, and hoping they don’t fire you. That’s essentially the argument more than 100 AI experts just made about the current state of AI safety evaluations.

The AI Evaluator Forum published a public letter on September 18 calling for genuinely independent third-party evaluators at frontier AI companies. The letter lays out a set of core conditions that its signatories say are non-negotiable for credible safety testing: independence, editorial control, transparency, protection from retaliation, and full access to the systems being evaluated.

What the letter actually says

The AEF’s letter is a direct response to proposals floated by two of the most powerful people in AI. Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman have both suggested a model where evaluators would be “nested” inside their companies, receiving “employee-like access” to test AI systems. On paper, that sounds reasonable. In practice, the AEF argues, it creates a structural conflict of interest that undermines the entire point of external evaluation.

The distinction matters. Employee-like access is not the same as employee-like dependence. The AEF wants evaluators to have deep technical access to models, training data, and internal documentation, but without the strings that come with being embedded in a company’s organizational hierarchy.

Conrad Stosz, the AEF’s chair, emphasized the importance of establishing shared principles that make independent oversight actually functional. The letter frames these conditions not as aspirational goals but as baseline requirements. Without them, the signatories warn, safety evaluations risk becoming a rubber stamp rather than a genuine check on increasingly powerful AI systems.

The five core conditions outlined in the letter form a clear framework. First, evaluators must be structurally independent from the companies they assess. Second, they need diverse viewpoints, not a homogenous team that shares the same assumptions as the developers. Third, transparency: findings should be published, not buried in internal memos. Fourth, protections against retaliation for evaluators who surface uncomfortable results. Fifth, full access to the systems and data needed to conduct meaningful testing.

The AEF-1 standard and its trajectory

This letter didn’t emerge from a vacuum. The AI Evaluator Forum launched its AEF-1 standard around December 2025, creating a formalized benchmark for how independent AI evaluations should be conducted. Since then, major companies including Anthropic and Google have begun implementing the standard, at least in part.

The fact that Anthropic appears on both sides of this conversation is telling. The company has engaged with AEF-1 implementation while its CEO simultaneously proposed a nested evaluator model that the AEF now explicitly pushes back against.

The broader context here is a regulatory landscape that’s still catching up. The AEF’s letter explicitly states that independent evaluations should support, not substitute, broader external oversight regimes.

Why this matters beyond the AI industry

The push for independent AI oversight carries real implications for how capital flows in the technology sector. Companies that adopt rigorous, verifiable evaluation practices position themselves favorably with institutional investors who increasingly factor governance and risk management into their allocation decisions. Firms that resist or delay adoption face a different calculus: potential reputational damage and regulatory penalties.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article