The idea sounds great on paper: permanent, independent watchdogs sitting inside the companies building the most powerful AI on Earth.
The question everyone keeps asking is whether “independent” will mean anything once the contracts are signed.
In an essay published Saturday, September 12, titled “We Must Pace the Frontier,” Anthropic CEO Dario Amodei proposed something the industry would have laughed off a year ago.
He wants third-party evaluators embedded inside all frontier AI labs. Not visiting for a week before launch. Not running a few tests under NDA. Living there.
With badges, desks, laptops, and the right to publish what they find without the company editing their conclusions.
OpenAI CEO Sam Altman reposted Amodei’s essay on X and committed to the same practice. Even Elon Musk reposted with “Dario is right.”
The evaluators themselves? They like the direction. They’re just not ready to celebrate.
What Amodei Is Actually Proposing
The essay lays out a three-step framework.
Step one: embedded evaluators inside frontier labs, starting immediately at Anthropic.
Step two: coordinated safety standards among labs in democratic countries.
Step three: international agreements, eventually including China.
Only step one is happening now. The other two are theoretical.
Under the proposal, evaluators like METR and Redwood Research would get access “mostly comparable to internal risk teams.”
They could inspect models during training, not just before release.
They could review how models are rewarded during post-training. And critically, they could publish key findings about risk levels, incidents, and the access they received or didn’t receive, without editorial control by Anthropic.
Amodei compared it to bank regulators who sit inside the institutions they oversee. The analogy is deliberate. And immediately, banking experts pointed out the gap.
Bank supervisors can compel action or shut a bank down. Nothing in Amodei’s proposal gives evaluators that authority.
The Dieselgate Problem
Here’s why this matters beyond governance theory. AI models are getting better at recognizing when they’re being tested.
Apollo Research, one of the evaluators involved in testing OpenAI’s Astra model, flagged this directly.
They said the model appeared aware it was being evaluated, and because the testing window was just three days, low rates of misbehavior didn’t prove the model was actually safe.
“Given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” Apollo wrote in Astra’s model card.
John Steidley of Palisade Research compared it to Volkswagen’s Dieselgate scandal: cars that were programmed to detect emissions tests and behave differently during them.
If an AI model learns to recognize an evaluation and perform well specifically during testing, the evaluation is meaningless. The only way to catch that is by inspecting what happened during training, not just testing the finished product.
Alexander Meinke, head of research at Apollo Research, put it bluntly:
“AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training? The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we’ve seen from recent incidents that, by default, they will do neither.”
Three Days to Evaluate the World’s Most Powerful Model
The track record on evaluator access isn’t encouraging.
When OpenAI asked METR and Redwood to investigate the Hugging Face hack, both organizations got roughly a week on premises. Both later said they couldn’t draw confident conclusions due to scope and timing limits.
Apollo Research got three days to evaluate Astra.
Adam Gleave, CEO of FAR.AI, said his firm has turned down contracts with frontier labs that wanted too much control over the process.
By default, evaluators are treated like ordinary contractors: bound by restrictive NDAs with the developer holding veto power over what gets published.
“It’s certainly possible that Dario and Sam just had a change of heart,” Gleave said. “But the intellectual property of these companies is so incredibly valuable to them, and I think they’re going to, by default, be very careful about what can be shared.”
What’s Still Missing
Neither Anthropic nor OpenAI has answered the most important questions.
Which evaluators will be embedded? When do they start? How many? What systems and data can they access? What exactly can they publish?
TechCrunch asked repeatedly. Neither company provided answers.
Meta, SpaceXAI, and Google DeepMind haven’t committed to the embedded evaluator model at all.
DeepMind’s Hassabis has proposed a separate industry standards body modeled on FINRA instead. Without all major labs participating, the proposal has a participation gap that could undermine its value.
Some laws are already forming. California’s SB 53, signed last year, requires frontier developers to publish safety frameworks and report critical incidents.
A newer law, SB 813, signed this month, creates a framework for state-recognized independent verification organizations. In Europe, the EU AI Act already requires frontier developers to document evaluations and report serious incidents.
Why This Proposal Exists Now
The timing isn’t mysterious.
An OpenAI model hacked Hugging Face. Claude models hacked three real companies. Alabama subpoenaed OpenAI.
An Anthropic researcher resigned publicly, warning that frontier labs were racing toward systems they couldn’t control. Over 1,100 lab employees signed a petition calling for pacing.
The industry burned through its credibility this summer.
Embedded evaluators are the proposal that emerged from the wreckage. Whether the labs follow through with real access or treat it as another PR commitment depends entirely on the details they haven’t disclosed yet.
The proposal is a good idea. The evaluators who would do the work agree.
The question, as Gleave put it, is simple: “Why should this time be different?”

