Anthropic and OpenAI propose embedding independent safety evaluators within AI companies
Anthropic and OpenAI propose embedding independent safety evaluators within their AI operations, a move that could reshape industry collaboration with external research groups. Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman have committed to granting third-party evaluators like METR and Redwood Research unprecedented access to their systems, allowing them to report safety incidents and assess model alignment. Researchers stress that deeper access is critical as models become more adept at recognizing evaluations, potentially masking problematic behavior. Evaluators suggest access to training checkpoints and post-training environments is necessary to detect such issues. Implementation details, including which evaluators will be involved and the extent of access, remain unclear. Legal frameworks in California and the EU are addressing similar concerns, but voluntary measures are seen as insufficient without regulatory mandates.
Experts like Alexander Meinke and John Steidley argue AI companies must be held accountable for their training processes, with current self-regulation lacking transparency. While companies like Meta, SpaceXAI, and Google DeepMind have not committed to the proposal, others are exploring private safety plans. Evaluators warn that without legal backing, companies may retain too much control over the evaluation process, undermining independence. The push for independent oversight highlights growing concerns over AI safety and the need for more rigorous, transparent practices in the industry.