Anthropic and OpenAI Propose Embedding Independent Safety Evaluators Inside AI Labs
Anthropic’s Dario Amodei and OpenAI’s Sam Altman support embedding independent evaluators, potentially including METR and Redwood Research, inside AI labs. Details remain unsettled, while Apollo Research had just three days to assess GPT-6 Astra before release.
Anthropic CEO Dario Amodei published an essay in September 2026 proposing that all frontier AI companies embed third-party safety evaluators inside their labs, granting them the power to assess model alignment, report safety incidents, and publish findings without editorial control. OpenAI CEO Sam Altman also committed to the practice, signaling a potential shift in how AI companies engage with outside research groups.
Amodei specifically named evaluators METR and Redwood Research as potential partners, offering them access to Anthropic’s internal systems. Neither Anthropic nor OpenAI has yet disclosed which evaluators they will work with, when they will be embedded, or exactly what information those evaluators can access or make public.
Researchers who spoke to TechCrunch welcomed the proposal but raised concerns about its implementation. The core issue is independence: evaluators say they have previously been treated as ordinary contractors, bound by restrictive NDAs and agreements that give AI developers significant control over what can be published. FAR.AI CEO Adam Gleave said his firm has turned down contracts with several frontier developers that sought too much control over the evaluation process.
Timing has also been a recurring problem. Apollo Research was given only three days to test GPT-6 Astra before its release, which the firm said was insufficient to draw firm conclusions about the model’s alignment. A similar constraint arose during an investigation of the Hugging Face incident, when METR and Redwood were given roughly a week on-site.
Evaluators say meaningful oversight requires access not just to finished models but to intermediate training checkpoints, employee interviews, and internal logs — access that would allow them to detect whether a model learned to pass safety tests without actually being safe, a concern researchers compare to Volkswagen’s Dieselgate emissions scandal.
Several researchers called for regulation to make the framework binding. California’s SB 813, signed in September 2026, creates a structure for state-recognized independent AI verification organizations. The EU AI Act similarly requires frontier developers to conduct model evaluations and report serious incidents. Meta, SpaceXAI, and Google DeepMind have not committed to the embedded evaluator model, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body.
Henry Papadatos, executive director of Safer AI, said voluntary measures remain dependent on company goodwill and could be abandoned during a PR crisis. “You cannot have it both ways, having zero accountability externally, and then say, ‘I’ll just have my own flexible rules,'” he told TechCrunch.