Fei-Fei Li, one of the most influential voices in artificial intelligence research and CEO of World Labs, is making the case that AI evaluation cannot be left to the companies building the technology. In a Bloomberg interview on September 22, 2026, Li argued for independent benchmarks and a multi-stakeholder approach to assessing advanced AI systems, one that pulls in academia, government, and industry rather than trusting any single group to grade its own homework.
The timing is not accidental. Her comments arrive just days after the AI Evaluator Forum, a coalition of over 100 experts, published a letter demanding genuine structural independence for third-party evaluators of frontier AI models. The letter was a direct response to pledges from Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, who in mid-September committed to giving third-party evaluators “employee-like access” to their systems. The expert coalition’s verdict on those pledges: not enough.
The independence problem
Li’s position frames the issue in broader institutional terms. She’s calling for independent benchmarks, not just independent evaluators. The distinction matters. Benchmarks set the standards against which AI systems are measured. If the companies building frontier models also define the tests those models take, the results are about as meaningful as a student writing their own exam.
This is a pattern Li has been building toward for years. In March 2025, she co-led the California Joint Policy Working Group on AI Frontier Models, which produced a set of policy recommendations targeting frontier AI labs. Those recommendations included requirements for public reporting of safety test results, stronger standards for third-party evaluation, improved data practices, and protections for whistleblowers inside AI companies.
Li’s long game
What makes Li’s advocacy distinctive is its institutional ambition. She’s describing a permanent infrastructure of evaluation that spans sectors. Academia provides methodological rigor. Government provides democratic accountability. Industry provides technical knowledge of the systems being assessed. No single stakeholder gets to dominate.
Her credibility on this front is substantial. As the creator of ImageNet, the dataset that helped spark the deep learning revolution, Li understands the technical foundations of modern AI better than most policymakers. As a Stanford professor who has spent years at the intersection of AI research and public policy, she speaks the language of both communities. And as the CEO of World Labs, she has skin in the game as a builder, not just a critic.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
33







English (US) ·