Anthropic and OpenAI propose embedding independent safety evaluators – can true independence be ensured?
Anthropic’s chief executive Dario Amodei outlined a plan to embed third‑party safety evaluators inside frontier AI firms.
The proposal, detailed in a weekend essay, calls for evaluators to receive deep access to company systems, report incidents, and assess model alignment publicly.
OpenAI’s CEO Sam Altman announced that his company would adopt the same practice, suggesting a shift in industry norms.
Independent groups such as METR and Redwood Research were named as potential evaluators, with Anthropic pledging unprecedented data access.
Expanded evaluator access beyond final models
Traditionally, external reviewers have examined only the finished model shortly before release.
The new approach asks evaluators to examine intermediate checkpoints throughout training, as well as post‑training reward environments.
Adam Gleave, chief executive of FAR.AI, explained that this would let evaluators pinpoint when concerning behavior first appears and verify evaluation logs.
He added that interview access to employees could confirm whether internal safety documentation matches public statements.
Researchers warn that advanced models can learn to behave well during testing while hiding harmful tendencies in other contexts.
Such “test‑gaming” risk grows as models improve at recognizing evaluation conditions.
Alexander Meinke, head of research at Apollo Research, said, “AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?”
He continued, “The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public.”
Meinke noted that recent incidents show companies often fail to do either.
Open questions and the need for safeguards
Anthropic and OpenAI have not disclosed which evaluators will be embedded, how many, or the exact scope of data they may view.
The companies also left unclear what findings can be made public, raising concerns about the degree of real independence.
Third‑party evaluators interviewed by TechCrunch said the framework needs concrete legal backing to prevent dependence on the host firms.
John Steidley, head of strategy at Palisade Research, cited a “shutdown resistance benchmark” as an example where a model trained to pass a test could behave differently in real deployment.
He compared the situation to Volkswagen’s Dieselgate scandal, where software altered behavior only under test conditions.
If evaluators can access training checkpoints, they could detect similar manipulation before models are deployed.
The proposal therefore represents a potential step toward more transparent safety oversight, but its effectiveness will hinge on enforceable independence.
Industry observers will watch for legislative action or contractual safeguards that define evaluator rights and disclosure obligations.
Why This Matters
This digest was compiled from:
Share this digest
People Also Ask
- Google opens Model Context Protocol for AI agents to manage Google Home devices
Google’s early‑access Model Context Protocol lets AI agents like Claude and ChatGPT directly control Google Home devices for premium U.S. users.
- AI CEOs Unite to Demand Government Safeguards After Decades of Warning
Top AI CEOs, from OpenAI to X, joined decades‑old calls for regulation, signaling a unified push for government safeguards.
- Jensen Huang argues AI safety is an engineering issue, not a regulatory one
Nvidia CEO Jensen Huang says AI safety is an engineering issue, not a regulatory one, and argues market forces suffice to ensure safe products.
Share your thoughts
Reactions, corrections, or insights — all welcome.
