The biggest AI risk is the system around the model.

published
15/09/2026
author
Mark Sheldon
read
5 min

An agent that can only draft a reply needs a reviewer. An agent that can call a buyer, qualify an invoice or change a collections strategy needs a production line around it. Every agent at Sidetrade runs on SAFE, the Sidetrade Agentic Framework for Enterprise, and when a CIO evaluates an agent that acts, the question underneath every other question is where the controls sit. This is how I answer it, layer by layer.

01 / the constraint

Governance as a factory, not a policy folder

Most governance programs are a folder: a policy, a committee, a review before launch. A folder does not stop anything, which is the whole argument for a factory. Each layer of the line has a defined job, a named owner and an inspection point, and a defect caught at intake never reaches the customer. Miss one layer and the risk travels down the entire line, because nothing downstream was designed to catch what that layer was supposed to catch.

The constraint that forces the factory is the action. Aimie calls buyers and qualifies invoices in production. A duplicate draft is noise. A duplicate call reaches a customer. For a finance function under SOX or the EU AI Act, an action nobody can reconstruct is unusable whatever its accuracy, so the controls have to exist as running code and written records, and a signed document is neither.

02 / the approach

The seven layers

Each layer below has a job and a diagnostic. The diagnostic is the inspection point: a question you can put to a team today and get a yes or a no.

01
Risk appetite, the blueprint

Write down what the AI may attempt, what it must refuse and what it must escalate, and get named owners to sign it before anything is built. Diagnostic: can every team state the prohibited actions in one minute?

02
Data quality, raw material intake

Check provenance, permission, completeness and freshness before data reaches the line. Nothing good gets built on bad stock. Diagnostic: would you know when an upstream source became unreliable?

03
Model inventory, the asset register

Track every model, version, use case, owner, vendor and deployment in one place. You cannot govern an asset you cannot see. Diagnostic: can one register show exactly what is running today?

04
Controls, line safety systems

Limits run automatically and stop the process when something is wrong. Thresholds, permissions, fallbacks and failure states are tested before launch, not discovered after it. Diagnostic: can a control fail closed with nobody watching a dashboard?

05
Human oversight, the floor manager

Define when a person intervenes and who owns the decisions a machine cannot make. A high-consequence action needs approval before it leaves the system. Diagnostic: does the reviewer get enough context to make the call?

06
Audit trail, quality records

Log inputs, model version, retrieved evidence, tool calls, approvals and outputs. The record has to prove what happened, explain why, and survive a regulator's questions. Diagnostic: could you reconstruct one disputed decision end to end?

07
Evals, final inspection

Test whether the system is correct, safe and useful before release, against explicit acceptance criteria for quality, cost and risk. Diagnostic: do evals rerun after every material model or workflow change?

The modelThe model is only one machine on the line. Governance is what turns an impressive demo into a controlled operating system.
03 / in production

How we run it in SAFE

{{author to confirm: one concrete example per layer from SAFE or an internal use case. Where the risk appetite is written and who signed it. How upstream data sources are checked for freshness. What the model register is and what it shows today. One control that fails closed, and what it stops. Where the approval threshold lives and who owns it. What one audit record contains. When evals rerun.}}

{{author to confirm: what this design builds on, if anything. NIST AI RMF, ISO 42001, the EU AI Act, or a specific framework or author.}}

04 / the numbers

The results

{{author to confirm: scope of the figures below, which use case or which period they cover.}}

Time to reconstruct a disputed decision

{{before}} to {{after}}

Agent actions that pass through an approval gate

{{share, before and after}}

Eval cycle after a model or workflow change

{{before}} to {{after}}

05 / the trade

What it costs

The factory is slower. Every layer adds an owner who has to sign and an eval that has to rerun before autonomy expands, and a blank at any layer blocks the expansion. I would take that trade again, because the alternative is finding the missing layer after the action has already left the system. {{author to confirm: which layer was hardest to stand up, or came in worse than hoped.}}

For founders, the first move is one real use case. Map all seven layers. Assign an owner and evidence to each. Fix every blank before expanding autonomy.

06

Questions this post answers

What is the biggest risk in deploying AI agents? The system around the model, not the model. An agent that takes real-world actions needs risk appetite, data checks, a model register, automatic controls, human approval, an audit trail and evals, each with an owner. A gap in any one of them lets risk travel to the customer.

How should a company govern AI agents that act autonomously? Treat governance as a production line rather than a policy folder. Give each of the seven layers a defined job, a named owner and an inspection point, and require approval before a high-consequence action leaves the system.

What should an AI audit trail contain? Inputs, model version, retrieved evidence, tool calls, approvals and outputs for every decision. The test is whether one disputed decision can be reconstructed end to end and explained to a regulator.

written by

CTO at Sidetrade. The founder of AI startup that was acquired by Sidetrade in 2016. 10 years immersed in AI & Agentic for B2B, focused on building Sovereign Intelligence.