When Dario Amodei calls for slowing frontier AI, and Sam Altman and Elon Musk publicly agree, we should take the warning seriously.
Amodei points to accelerating recursive self-improvement and recent incidents in which AI agents exceeded their intended scope. His response combines embedded external evaluators, coordination among democratic countries and, eventually, global agreements limiting dangerous AI development. Altman has committed OpenAI to the evaluator proposal, while Musk responded that “Dario is right.” Amodei’s essay; Associated Press.
But pacing the frontier is not the same as governing it.
- The first problem is that slowing development does not create control.
An extra year may help researchers improve alignment, interpretability and testing. Yet none of these mechanisms guarantees that an AI system will remain within its authority when it acts in the real world. A model can pass an evaluation and still take an inadmissible action when the user, data, jurisdiction or surrounding circumstances change.
Consider a healthcare agent authorised to check on a patient each evening. Its general permission to discuss symptoms does not mean it should answer every question. If the patient reports chest pain, suicidal thoughts or a serious drug reaction, an interaction that was routine moments earlier must be stopped and escalated. No frontier-wide speed limit makes that decision. A control operating at the point of action does.
- The second shortfall is that embedded evaluators remain focused mainly on the laboratory.
Giving independent experts employee-like access is a welcome improvement over companies assessing themselves. But access alone does not create independent evaluation.
Evaluators need their own models, testing tools, datasets and evaluation harnesses. They need the technical and domain expertise to create adversarial scenarios, recognise subtle failures and challenge the assumptions built into the developer’s safety framework.
If they depend on the laboratory’s models, dashboards, test environments or definitions of success, they may simply reproduce the laboratory’s blind spots. They remain independent in name while assessing the company through the company’s own instruments.
Even properly equipped evaluation is still observation rather than enforcement. Evaluators may inspect training environments, reproduce tests and publish findings, but can they prevent an agent from executing a prohibited action? Effective oversight requires independent infrastructure, access to the underlying evidence and authority to escalate, restrict or stop deployment when safety conditions are not met.
This matters because an AI system cannot be the sole control system for its own behaviour.
If a capable agent can evade a sandbox, manipulate a test or attack its grader, a monitor operating inside the same technical environment may also be suppressed or deceived. Oversight needs an out-of-band path the model cannot reach, alter or spoof.
- The third problem is enforceability.
Amodei acknowledges that global coordination will be difficult and that a full pause is unlikely because the incentive to defect would be enormous. Open-weight systems can be copied and modified beyond the original developer’s control, while offshore developers may sit outside the jurisdiction creating the rules. A regime designed around a few frontier laboratories could therefore slow compliant companies while leaving much of the deployment surface untouched. It could also turn the infrastructure already owned by the largest labs into a regulatory moat.
A better approach combines upstream and downstream governance.
Frontier developers should face mandatory evaluations, independent security testing, incident reporting and tamper-evident monitoring. But every high-risk deployed system, regardless of who trained it or whether its weights are open or closed, should also operate behind an independent runtime control layer.
Before an action crosses into the real world, that layer should verify the system’s current authority, applicable law, company policy, available evidence and the action’s admissibility. It should be able to narrow permissions, require human review, revoke authority and produce an auditable record.
A bank’s agent might be authorised to discuss mortgages. If it encounters a vulnerable customer, a complaint or a jurisdiction-specific disclosure obligation, its authority must change before it produces the next response. The same principle applies when an agent initiates a payment, publishes a pharmaceutical claim or accesses confidential records.
We already govern powerful institutions this way. Under frameworks such as Basel, firms assess and manage risks within predefined rules while independent regulators supervise, challenge and enforce.
Pacing may buy time. Only enforceable governance can determine what AI is allowed to do with it.