Agency Is Not a Metaphor
When an AI agent acts on your behalf, a centuries-old legal doctrine already decides who's liable. The audit trail you skipped is the evidence that decides it.
We keep arguing about whether AI agents can be trusted to do real work. It’s the wrong argument. The law settled the interesting question a long time ago — we just haven’t noticed it applies.
“Agent” is a legal word before it is an AI word. Under the doctrine of agency — centuries old, boringly well-established — a principal is liable for the acts of an agent carried out within the scope of the authority the principal granted. Hire a broker and you answer for what the broker does in your name. The agent acts; the principal pays.
Now read “AI agent” again. The moment a system acts on your behalf — sends the message, books the appointment, approves the refund, files the claim — you are not the operator of a clever tool. You are the principal of an agent. And the question that follows you home is not “is the model good?” It is the old one: did it act within the authority I granted, and can I prove it?
It helps to picture how this actually gets decided, because it doesn’t happen in a press release. It gets reconstructed, after something has already gone wrong, by the people whose entire job is to work out what happened and who pays — adjusters, underwriters, regulators, opposing counsel. They are handed the aftermath and asked a deceptively simple question: at the moment of the act, what was the system allowed to do, what did it actually do, and where is the proof? Either that record exists, or it doesn’t. And when it doesn’t, the absence is not neutral — it is read against the party who should have kept it. The principal.
That is why the audit trail is not compliance theatre, and not IT hygiene. It is the evidence of granted authority and scope-adherence — the precise thing that decides whether the principal is on the hook. I wrote separately about the five artifacts a carrier actually asks for in an AI breach claim; this is why they ask. Four of the five are missing from most pipelines because nobody built them as liability evidence. That is exactly what they are.
The cleaner industries have already worked out where this lands. When carmakers began putting genuinely autonomous systems on the road, the serious ones stopped hiding behind the driver: Volvo and Mercedes each said, in effect, that when our system is in control, we accept the liability. It’s our product; we back it. One named principal, owning the outcome, on purpose.
Most companies will not get to make that choice as gracefully as Volvo did — the courts are already making it for them. When Air Canada’s chatbot invented a refund policy the airline then refused to honour, Air Canada argued the bot was a separate entity responsible for its own statements; the tribunal disagreed and held the airline to what the bot said, exactly as if a human agent had said it. This year a German court went further, ruling that an AI’s summaries are not a neutral relay of other people’s words but the company’s own speech — an expression of its business — so the company answers for the ones that mislead or defame, with an AI-defamation suit over a false summary already winding through the courts. The pattern repeats: try to disown the agent, and the disowning is what gets read against you.
AI agents have no single Volvo. There is the company that made the model, the company that deployed and orchestrated it, and the business that put it to work. When the agent acts and it goes wrong, three parties can each point at the other two. That diffusion feels comfortable — right until the first real claim, at which point “it’s complicated” is not an answer anyone wants to give under oath.
Someone has to be the principal who backs it. And the only party who can credibly do that is the one who can actually see and control the scope: what the agent was allowed to do, what it did, and the gate a human or a policy put in front of the consequential actions. That is the operator who runs the governed layer — not the model vendor, who cannot see your context, and not the busy customer, who never signed up to be the operator. Backing it is a choice. The evidence trail is what makes that choice survivable.
There is a quieter trap underneath all of this. The model is not a fixed thing you certified once. It versions. APIs iterate, weights update, a provider deprecates the endpoint you built on, and the behaviour you carefully validated last quarter quietly shifts — the thing that just finished baking is half-broken again, and nobody sent a memo. Now reconstruct liability: when the model silently changed under you, whose fault is the new failure?
You cannot answer that without pinning the version — which model, which version, which configuration, which prompt template, recorded against the specific act. That is not fastidiousness; it is the difference between “we can show precisely what produced this decision” and “we think it was probably the model.” A governed layer that pins the version, evaluates every change before it ships, and keeps the record of the drift is the only thing that turns a moving floor into something you can stand on in front of an adjuster.
Where regulators have already named who is accountable, this stops being abstract fast. UK financial services puts specified senior managers personally on the hook for what happens under them — a named human, not a department. The moment accountability has a name, the audit trail stops being an IT line item and becomes that person’s personal cover. Clarified accountability does not relieve the pressure to govern; it concentrates it on someone who now has an intensely personal reason to prove the agent behaved. That person is the real buyer of provable governance — and they sit well above procurement.
So here is the whole thing, stripped down. Liability follows authority. Authority has to be granted deliberately, bounded carefully, and — when the agent acts — provable after the fact. The party who can prove it is the party who governs it. Everyone else is hoping.
The agents can do the work; that argument is over. The unsettled question is who is the principal when they do — and whether that principal can show what was allowed, what happened, and where the human stood. “Governed” stops being a feature you weigh against price the moment it is the evidence keeping a named person out of the claim.
Pick your vendor on who is willing to be the principal.
