← Back to blog
AI governance

Could You Replace Your AI Model Without Rebuilding Governance

Fahd Rachidy, CEO ·

CIOs spend considerable time discussing AI vendor risk, model performance and the danger of technological lock-in. But there is a more revealing question: if you replaced the model underneath an AI agent tomorrow, could you preserve the same permissions, controls, escalation paths and audit evidence?

If the answer is no, the organization has outsourced part of its control environment to the model provider.

This is the model-swap test. It asks whether the organization’s governance survives when the intelligence inside the system changes, whether the autonomous system is safer or not, or uses open or closed weights.

That question matters because enterprises will change models frequently: for cost, performance, data residency, resilience, acquisition or simply because a better model becomes available.

The intelligence may change but the responsibility does not.

Take for example a bank using an agent to prepare customer communications and suitability reports. The bank may initially deploy a closed commercial model, then move part of the workflow to another provider or a privately hosted model.

The new model may interpret instructions differently. It may call tools in a different order, escalate fewer borderline cases or produce more confident language. Its reasoning style may change even though the application, customer data and business objective remain the same.

The bank’s obligations, however, have not changed.

It must still establish what the agent was permitted to do, whether the recommendation was supported by the customer’s circumstances, which policies applied and who approved any exception. And keep the full audit trail of any action.

And given that a model provider can now barely explain how its system was trained and evaluated, it cannot of course own the deploying institution’s legal duties, customer relationships or liability. Those remain with the enterprise before and after the model swap.

This distinction is becoming more important as models become harder to monitor internally.

We are entering a post-chain-of-thought world.

OpenAI Chief Scientist Jakub Pachocki recently wrote that the company’s ability to rely on chain-of-thought monitoring is “progressively diminishing.” Models are becoming better at reasoning about their own reasoning, blending internal thought with tool use and solving problems without verbalizing every meaningful step. He expects AI progress to become increasingly limited by confidence in monitoring. OpenAI’s “An Alien Mind”.

That is now directly becoming a CIO problem.

Indeed, many enterprise governance approaches assume that a model can explain why it made a decision. Organizations preserve the prompt, response and perhaps a model-generated rationale, then treat that record as evidence (although it is now known that prompts are not accepted as ‘legal’ evidence). But an explanation generated by the same system that took the action is still a self-report. It may be useful, but it is not an independent control and may not be a complete account of what drove the result.

Furthermore, as internal reasoning becomes less and less visible, CIOs need another way to prove that an agent remained under control. The answer might likely reside in preserving evidence outside it.

Three things every CIO must be able to prove.

These do not materially differ from the three lines of defense.

The first is authority. What was the agent allowed to do at that particular moment?

Authorization cannot be a permanent permission attached when the agent is deployed. An agent permitted to draft an email is not necessarily permitted to send it. An agent allowed to prepare a payment is not necessarily allowed to approve it. Authority may change when a customer becomes vulnerable, a transaction exceeds a threshold, data crosses a jurisdiction or a human withdraws consent.

The second is evidence. What information, rules and system state supported the action?

A useful record must include more than the conversation. It should identify the model and prompt versions, the proposed tool call, the relevant customer or transaction facts, the policy version applied, the decision made by the control layer and its citations and explanations of the decision, and what the external system ultimately executed. Otherwise, an audit log can show what the agent said without proving why the action was permissible.

The third is accountability. Who owned the outcome and who approved the exception?

Human oversight cannot mean placing a person somewhere near every single workflow. The system must record when review was required, what the reviewer saw, what they decided and whether the eventual action matched that approval. The business remains the first owner of the risk, compliance defines and oversees the applicable rules and autonomous authorization layer, and internal audit needs independent evidence that both performed their roles.

These three proofs must remain available even when the model cannot explain itself and even when the organization replaces that model entirely.

Separate model intelligence from enterprise authority.

Passing the model-swap test requires a deliberate architectural boundary.

The model should propose an action. An independent control should determine whether that action may be executed. The enforcement point should sit in the agent harness, where it can intercept consequential tool calls before they reach customer, financial, clinical or operational systems.

That control should evaluate the proposed action against current laws, regulations, authority, company policies, applicable obligations or industry standards, and the facts available at that moment. It must be able to allow the action, narrow it, call and require human review when needed, or simply block it. The resulting evidence should be written to a record the agent cannot alter.

This arrangement also changes how organizations evaluate models.

Testing cannot depend entirely on a provider’s model, datasets, definitions of success or evaluation harness. Enterprises and their independent evaluators need suitable tools, domain expertise and test environments that can compare providers against the same real business scenarios. Otherwise, each model swap changes both the system being assessed and the instrument used to assess it.

Portability therefore means more than placing a common API in front of several models. Technical abstraction may make models interchangeable, but governance portability requires consistent authority, evaluations, enforcement and evidence across them.

Put the model-swap test into procurement.

Before approving an AI platform, CIOs should ask whether the organization can export its operational evidence, pin model versions, detect provider changes and test a replacement model against the same controls. Contracts should define notification requirements for material model changes and preserve access to the information needed for investigations and audits.

Architecture reviews should then run the practical test: replace the model in a controlled environment and observe what breaks. If permissions, policy decisions or evidence disappear with the original provider, the organization has identified a governance dependency that ordinary vendor-risk questionnaires will miss.

The goal is to make every model operate within the same enterprise authority. Models will improve, become less expensive and be replaced. Their internal reasoning may also become increasingly difficult to inspect. CIOs should design for that reality now.

The model can be powerful and replaceable. Governance must be independent and durable.