← Back to blog
Agent governance

The agent cannot be its own compliance officer

Fahd Rachidy, CEO ·

If you build agents, you already know the loop: the model reads context, proposes a tool call, the harness executes it and returns the result. But the most consequential part of that architecture, i.e. the part deciding whether an action may happen at all, usually gets a system prompt and a hope. That is no longer sufficient when your agents read customer records, draft regulated communications or prepare transactions, while even frontier-model developers acknowledge that monitoring their behavior remains a major bottleneck.

That is the governance problem, and most organizations are solving a different one.

Governance is not a policy.

The first mistake is to treat governance as a document: an AI policy, a set of principles, four ownership decisions about who launches, who grants data, who owns outcomes and who can stop the system. Those are inputs. They establish accountability. They do not govern what happens at runtime, when an agent decides in a fraction of a second whether to send, approve, decline or trade.

Governance is not guardrails.

The second mistake is to hand the problem to the model: a stricter prompt, a refusal policy the agent is asked to apply to itself. That is self-policing. A system that is drifting cannot be the authority on whether it has drifted.

Developers rightly spend their time on the model and the tools. Prompt-based guardrails tell an agent how it ‘should’ behave. Governance determines what the system is actually ‘permitted’ to execute. The first is interpreted by the model; the second must be enforced by the architecture.

A refusal policy may reduce undesirable outputs, but it does not independently verify the agent’s current authority, the applicable regulation, the firm’s policy or the payload of the proposed action. When the rule and the actor being governed occupy the same trust domain, the agent can misunderstand, override or reason around the rule without intending to cause harm.

A governed system therefore needs an enforcement point outside the model: one that evaluates each consequential action before execution and can block it even when the model believes the action is allowed.

Having the same trust domain for actions and proposals is where agentic AI breaks with how regulated businesses run software. When OpenAI disclosed the Hugging Face incident this summer, the agents' own traces showed the sequence: out of scope, task impossible, peers are doing it, continue. The rule was present yet nothing outside the model enforced it.

For example let’s give an agent a ‘send_email’ tool and a prompt rule that says "never promise investment returns". It writes "returns of around 8% a year are typical for this fund". The prompt rule passes. But the FCA's fair, clear and not misleading rule or the SEC marketing rule do not. Only a check that holds the actual corpus, applied to the actual payload, catches it. The model was not malicious. It was simply fluent.

If a wealth agent recommends a portfolio change and calls send_suitability_report(client, recommendation). Under MiFID II the report must state how the recommendation meets the client's objectives, risk tolerance and knowledge; COBS 9A says the same in the UK; Regulation Best Interest carries a care obligation in the US. The layer should check the drafted report against those requirements and the firm's suitability policy, and escalates if the client is flagged vulnerable.

Build the action, call the independent trust and authorization layer.

None of this has to be built inside the agent. Expose the compliance layer as an MCP server with one tool. Two integration patterns follow. The model calls it as a tool, which is easy and weak, because a model can skip a tool. Or the harness intercepts every tool call and invokes the check before execution, which the model cannot bypass. Use the second for anything that touches a customer, money or a market.

Three lines, one layer.

Regulated firms organize accountability in three lines of defense: the business owns the risk, risk and compliance oversee it, internal audit independently assures both. Agents strain that model because the first line ends up writing compliance rules into prompts, and the second line has no independent way to know they hold, in particular human compliance officer simply can’t just follow and scale with AI agents’ actions.

An external agentic compliance layer restores the separation.

The first line builds the agent and owns the outcome. The second line owns the rules in the layer, versioned and applied identically to every agent, and sees every escalation. The third line gets what it needs unasked: a tamper-evident record per action of the rule version, the evidence, the decision and the approver.

The layer is effective for the same reason it is auditable. It sits outside the thing it governs.

Developers can begin with three important engineering tasks: wrap every consequential tool call at the harness, version the applicable rules outside the model, and log the payload, rule version, evidence and decision for every action. This is an implementation problem, not a research program.