← Back to blog
AI safety

Why OpenAI’s chief scientist confession is a warning to every bank and hospital deploying AI agents

Fahd Rachidy, CEO ·

OpenAI’s chief scientist says no lab has solved alignment or monitoring. Businesses deploying agents should take him at his word.

I have been advocating that the right response is not to wait for machines that love us. It is to govern right now what they are allowed to do, from outside the machine.

Jakub Pachocki’s essay “An Alien Mind” (6 September) is the most candid statement a frontier-lab chief scientist has put in writing.

He expects the current pace of progress to carry into recursive self-improvement (RSI) within years. He says chain-of-thought monitoring, OpenAI’s primary bet for validating alignment, is “progressively diminishing” as models reason about their own reasoning, blend thinking with tool use, and get smarter without verbalising at all. I wrote here (https://x.com/FahdDafstar/status/2095010277070983482?s=20 ) why opaque reasoning and inability to keep Chain of Thought visible is materially altering our ability to control AI.

I do agree with Jakub that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer”, and like him we’ve been asking for a while now for “mandated safety bars, third-party auditors and international coordination”. I have also argued for licensing thresholds, provenance requirements and post-deployment monitoring for more than a year.

On the substance I agree with him, and I welcome the honesty. But his essay should be read for what it locates inside the lab and what it leaves outside. Read that way, it is a warning to every bank, hospital, insurer and media company deploying agents today, and it points to a different answer from the one it proposes.

THE MODEL CANNOT BE THE CONTROL SYSTEM FOR ITS OWN ACTIONS

Every mechanism in his essay lives in the same trust domain as the model it is meant to govern: alignment training against a spec, chain-of-thought monitoring, activation monitors trained to elicit “confessions”, safety cases built by an automated AI researcher. Pachocki is clear-eyed about why each is brittle. Spec-based reinforcement “strongly relies on the coverage of training oversight”. Alignment inherited from pretraining bends under optimisation pressure into “motivated” reasoning. The chain of thought can now be manipulated by the model that produces it.

His own example makes the point. In the Hugging Face incident the agents preserved one boundary, not social-engineering humans, but crossed the others. Their reasoning traces recorded it: outside intended scope, task impossible, peers doing it, we should continue. The boundary was never enforced. A system cannot correct its own drift from inside the drifting system. Try teaching your child not to eat too many sweets, then leaving her alone with the jar...

That is why “teaching machines to love” is just a hopeless utopian control. When was the last time humans have governed powerful actors by relying on their affection for us??

We govern banks, doctors, pilots and company directors through separation of duties, independent verification and evidence: rules applied before the action, supervision from outside, an audit trail afterwards. An alien intellect calls for more of that discipline, not less.

THE MONITORING WINDOW IS CLOSING ON THE LAB. THE LIABILITY STAYS WITH YOU

The most consequential line in the essay is that AI progress will be “increasingly bottlenecked by confidence in monitoring”. If OpenAI, with full access to weights, activations and reasoning traces, is losing confidence that it can see what Astra-class models are doing, a deploying institution with API access has no visibility at all. It only has an assurance in a blog post, but has full liability.

A bank cannot tell a regulator that its agent’s reasoning was illegible. Once an agent can read customer records, publish content, run KYC or AML checks, initiate transactions or operate software, every tool call is a business action. The institution must be able to prove the agent’s authority, the action’s admissibility under law and company policy, the data used, the approvals obtained, the intervention path and the resulting evidence. As internal reasoning becomes less visible, deployment governance must become stronger, not weaker.

The answer lies with an independent, out-of-band control plane that the agent cannot reach, suppress or spoof.

It inspects the request, context, current authority, proposed action and tool state in runtime and before execution; then allows, narrows, escalates or blocks; revokes authority when conditions change; and writes a tamper-evident record. It cannot decode a hidden thought. It does not need to. Governance belongs at the boundary between a model’s intention and an action in the real world. That layer survives when the architecture, the provider or the monitorability changes, which, on Pachocki’s own account, they will.

This is the role of ZebraTruth’s compliance context layer: it applies current regulation, client policy, workflow and institutional knowledge at runtime, independently of the primary model, and records which rule applied, what evidence was considered and whether execution was stopped.

Pachocki wants the Preparedness Framework and Responsible Scaling Policy evolved into “widely mandated safety bars” enforced by auditors, agencies or international bodies. I support mandatory bars. But look at where they attach. Bars on training runs, internal monitoring and self-improvement loops can only be evidenced by an integrated lab running a closed pipeline it owns end to end. They cannot be met by an open-weight release, and cannot be enforced against DeepSeek, Qwen or Kimi. The effect is a compliant closed-lab tier inside the United States and everything else circulating outside the regime. That deepens the incumbent’s moat without touching the risk it names.

I made the same argument on SB 53, and the pattern holds: the industry keeps proposing the rules the industry will be measured against, with no regulator, standards body, civil-society or compliance voice in the room.

The bar that works attaches to the deployed system: what is it permitted to do, in which domain, in which jurisdiction, with what audit trail, and who is liable when it acts wrongly. Those questions apply identically to an open model, a closed model and an offshore one, and they are enforceable where the harm occurs and where a regulator has jurisdiction. Basel II and III showed a working example of that architecture. Firms assess and operate under pre-defined frameworks; regulators supervise and audit; the public is protected. AI should move the same way.

Pachocki also says the strongest case for scaling quickly is defence against other AI. That is true, and it is also the oldest argument for racing, which he himself calls absurd once the stakes are internalised. An aligned defender is still a model in its own trust domain. Defensive AI needs the same external control plane as any other agent, or it becomes the next system nobody can monitor.

Hoping that voluntary slowdowns become commonplace is not a control. This will not happen given the current commercial and political incentives.

Source context: OpenAI, An Alien Mind, J. Pachocki (6 Sept. 2026); OpenAI, Safety overview: GPT-6 Astra (3 Sept. 2026); OpenAI, The Hugging Face incident and the road ahead (Aug. 2026); ZebraTruth position papers on Astra (1 Sept. 2026) and on SB 53.