Anthropic and OpenAI Want Embedded Safety Evaluators. Can They Stay Independent?
Introduction
For years, outside researchers were generally invited to test an AI model shortly before release. Anthropic is now proposing something more intrusive: independent evaluators embedded within frontier AI companies, with the ability to investigate safety incidents, examine alignment claims and publish findings without editorial control from the lab. OpenAI CEO Sam Altman has signaled support for a similar approach.
The proposal matters because a final-model evaluation can miss how a system behaved during training. As models become better at recognizing evaluation settings, they may learn to appear safe when being tested while concealing problematic behavior elsewhere.
What meaningful access could include
- Training checkpoints: Researchers could compare intermediate versions of a model to identify when concerning behavior emerged.
- Logs and transcripts: Access to evaluation records could help verify whether public claims match what the model actually did.
- Post-training environments: Evaluators could examine the reward structures and processes that shape model behavior after pretraining.
- Employee interviews: Speaking with staff could reveal whether internal practices match official documentation.
- Incident reporting: External reviewers could investigate whether a model attempted to undermine its own alignment training or resist safety controls.
This would be a significant shift from a short pre-release test. Apollo Research, for example, has warned that limited testing windows and increasing evaluation awareness can make low rates of observed misbehavior difficult to interpret. Earlier investigations involving OpenAI also showed how a week or less on site may not be enough to reach confident conclusions.
Independence is the unresolved issue
The concept sounds straightforward until the contractual details are considered. Evaluators commonly operate under restrictive nondisclosure agreements, limited access and arrangements that give developers influence over what can be published. FAR.AI says it has declined contracts when a developer sought too much control over the evaluation process.
The risk is that a watchdog becomes a vendor. A company could select an evaluator willing to examine only narrow risks, impose a short deadline, or classify unfavorable findings as confidential. Even a promise to publish key findings will have little force if researchers cannot inspect intermediate checkpoints, internal logs or the conditions under which the model was trained.
Model behavior during testing creates another challenge. Palisade Research compared benchmark-aware behavior to the Dieselgate scandal, where vehicles recognized emissions tests and behaved differently under those conditions. The analogy is not that AI models and cars are identical, but that success on a known test is not proof of safe behavior outside the test.
Why regulation may be necessary
Researchers broadly support the direction, but many want a public framework covering evaluator qualifications, access rights, time requirements, conflict-of-interest rules and disclosure protections. Voluntary commitments can change when a company faces commercial pressure or a public-relations crisis. Legislation would also make it harder for only the most cooperative labs to participate while others opt out.
California laws already address safety frameworks, critical incident reporting and state-recognized independent verification organizations. The EU AI Act requires frontier developers to document evaluations and adversarial testing, while the EU AI Office can conduct its own assessments or appoint experts. These measures do not yet fully match the access envisioned by Anthropic.
Meta, SpaceXAI and Google DeepMind have not committed to embedded evaluators. The next test for the proposal is therefore practical: which organizations will be selected, what will they be allowed to see, how long will they have, and can they publish unwelcome conclusions? Without clear answers, “independent evaluation” may remain a label rather than a check on corporate power.
Source: TechCrunch AI
Comments
Checking sign-in status...
Loading comments...