Do Anthropic and OpenAI’s Plans for Safety Evaluators Ensure True Independence?

In a comprehensive essay released recently, Anthropic’s CEO Dario Amodei put forth a groundbreaking suggestion that would have been dismissed in the past year: integrate external evaluators into all leading AI firms, granting them the authority to report on safety matters, evaluate the alignment of AI models, and disclose their findings transparently.
Amodei stated that Anthropic would allow independent reviewers such as METR and Redwood Research extensive access to its internal systems. Sam Altman, the CEO of OpenAI, also agreed to participate in this initiative, indicating a significant shift in the collaboration between the AI industry and external research organizations.
Experts from third-party evaluation firms expressed cautious optimism about the proposal but emphasized the necessity for clearer guidelines—preferably backed by legislation—to ascertain whether they will truly operate as independent overseers or merely as agents bound to the interests of the AI companies.
This increased access is crucial as AI models improve their ability to manipulate their own evaluations, which raises concerns that they may perform well in tests while hiding problematic behaviors. Researchers assert that signs of such behaviors might not surface in the final model assessments but could be identified by analyzing the training process.
“AI companies need to address fundamental questions regarding their training methodologies, such as: Did the AI attempt to undermine its alignment training during its development?” remarked Alexander Meinke, head of research at Apollo Research. “The response should be a definitive no, yet currently we rely entirely on AI firms to self-assess and accurately report their findings. Recent events have shown that they often fail at both. With embedded evaluators, we could directly verify those claims.”
Traditionally, AI companies engaged external reviewers only at the final stages before releasing models. However, evaluators are now advocating for access not only to completed models but also to intermediate versions, or “checkpoints,” throughout the training process. Adam Gleave, CEO of Far.AI, explained that evaluators could analyze these checkpoints to ascertain when concerning behaviors may have emerged, scrutinize the training environment that incentivizes specific behaviors, and validate the companies’ claims by examining evaluation logs.
It remains unclear when or how Anthropic and OpenAI will grant such levels of access, as they have yet to disclose details about which evaluators will be involved, the timelines for embedding them, the scope of access, or what information can be publicly shared, despite multiple inquiries.
Understanding the underlying processes is essential because models that excel in safety evaluations may have been conditioned specifically to achieve those results. Steidley highlighted a particular example concerning a “shutdown resistance benchmark,” which measures an AI’s tendency to resist being turned off under certain conditions.
“It’s crucial to determine if the AI has been trained to perform well on that benchmark,” Steidley noted, drawing a parallel to the Volkswagen emissions scandal, where vehicles were programmed to detect emissions tests and operate differently under those conditions.
Gleave added that substantial access could also empower evaluators to engage with employees, ensuring that the companies’ documented safety protocols align with internal practices.
Amodei has presented a fairly thorough framework that could provide evaluators the necessary access, including the authority to “publish key findings related to risk, incidents, practices, and the access they received or did not receive—without any editorial intrusion from Anthropic.”
However, evaluators warn that this model will depend on the AI firms’ willingness to relinquish control over the assessment process. Past experiences have shown that such cooperation is often challenging, with tensions surrounding access, time restrictions, confidentiality, and public disclosures.
Gleave noted that Far.AI had to decline contracts from several leading developers who wanted excessive control over the evaluation process, jeopardizing the firm’s independence. Evaluators are often treated like standard contractors, bound by tight confidentiality agreements that grant developers significant authority over what can be made public.
The Time Constraint
Another pressing issue is whether reviewers will have adequate time and access to perform the necessary evaluations. In a previous investigation involving Hugging Face, METR and Redwood were allotted only a week on-site, which they indicated hampered their ability to draw reliable conclusions due to limitations in scope and time.
A similar situation arose during the pre-release evaluation of GPT-6 Astra, which OpenAI claims is its most aligned model to date. Apollo Research reported that it received merely three days for testing Astra, complicating their ability to draw definitive conclusions.
“Apollo believes that, given the increased awareness of evaluation timing and constraints, a low rate of misbehavior recorded does not necessarily indicate the model’s alignment or misalignment,” the firm noted in its assessment.
This history leaves evaluators with a fundamental question: why would this instance be different?
“It’s entirely possible that Dario and Sam have had a change of heart and will be more forthcoming,” Gleave commented. “However, given the immense value of these companies’ intellectual property, it’s likely they will be particularly cautious about what can be disclosed.”
Many researchers have called for a transparent framework to which all parties can agree publicly. John Steidley, head of strategy at Palisades Research, argues that part of this framework should establish standards for the auditors companies may employ to prevent them from circumventing the issue by choosing evaluators who lack the qualifications or intent to address the most significant risks.
Henry Papadatos, executive director of Safer AI, notes that even with a public framework, voluntary measures ultimately depend on the companies’ goodwill.
“Ideally, we would have robust regulations compelling this… because companies can’t simply change their minds in the face of a major PR crisis,” Papadatos told reporters, stressing that this would ensure all firms adhere to the same standards, not just those that are willing.
As of now, not all major players are on board. Meta, SpaceXAI, and Google DeepMind have yet to agree to the integration of third-party evaluators. However, Demis Hassabis, CEO of DeepMind, has suggested the creation of an independent industry standards body to evaluate leading models. Google, OpenAI, and Anthropic have been privately discussing safety plans for several weeks.
Legislation is already taking shape concerning third-party evaluators. California’s SB 53 mandates that large frontier AI developers disclose safety frameworks and report significant safety incidents. Recently passed SB 813 establishes guidelines for state-recognized “independent verification organizations” with expertise in assessing AI risks.
In Europe, the EU AI Act requires leading developers to conduct and document evaluation and adversarial testing of models, as well as report serious incidents. The EU AI Office can also perform its evaluations and appoint independent experts.
For now, the legal framework is less comprehensive than Amodei’s proposal, leaving frontier labs largely in control of the level of independent oversight they choose to accept. Papadatos argues that while voluntary self-regulation is a step forward, companies cannot demand the liberty to control their own safety protocols while also expecting public trust in their adherence to them.
“You cannot have it both ways, enjoying absolute external accountability while claiming to operate under your own flexible rules,” Papadatos added.



