Who answers when an agent answers

The question that decides whether an agent reaches production isn’t technical. It’s who owns the outcome when it goes wrong.

Carlos Andrés Ramírez ·

Almost every conversation about agents starts with capability. What the model can do, how far it goes, how good it is. Then it gets stuck somewhere else entirely: what happens the day it does something nobody wanted.

The pattern repeats often enough to stop treating it as coincidence. A team builds an agent that works. They demo it. Everyone nods. And then the project enters a phase nobody planned for, which can last months, in which the agent is no longer improved: it’s debated.

What gets debated is never accuracy. It’s a much simpler and far more uncomfortable question: if this agent approves a discount it shouldn’t have, whose mistake is it?

Two different questions get mixed up here. Who is legally liable when an agent causes harm gets settled by a court or a contract, case by case, and law firms already specialise in exactly that. Who answers inside the company, every day, before anything goes wrong, is a separate question, and nobody handed it to outside counsel or to the platform vendor. Buying an agent with an audit trail doesn’t settle it either: the trail logs what the agent decided. It never says who decided the agent could decide that on its own.

The symptom

Accountability doesn’t delegate upwards.

When an organisation hasn’t decided who owns the outcome, the default answer is to escalate it. The agent proposes and a person confirms. On paper that sounds prudent. In practice it’s the most expensive way to automate nothing: you’ve built a system that does the work twice, once by the agent and once by whoever reviews it, and added a queue in between.

  • The agent proposes, but nobody has the authority to stop reviewing it.
  • The review isn’t measured, so nobody knows how often it actually changes anything.
  • Every exception moves up a level, and that level lacks the context to decide it.
  • The team that built the agent isn’t the team answering for its decisions.
  • There’s no written threshold above which the agent decides on its own.

The problem underneath

An agent isn’t a tool. It holds a post.

A tool is used by someone, and that someone answers for the result. An agent acts on the organisation’s behalf inside a process, which puts it far closer to a job than to a piece of software. Jobs have a decision scope, a manager, a threshold above which they escalate, and a measure of whether they’re doing well. Agents that reach production have all four. Agents that stay in pilot have none.

If you can’t name the person who answers for what the agent decides, you don’t have an agent in production: you have a demo with an on-call rota.

BECOME

What changes

Four decisions, before the first line of code.

Scope
Which decisions sit inside the agent and which don’t. Written as a closed list, not as a general principle.
Owner
A named person who answers for the aggregate outcome. Not the AI team: whoever already answers for that process.
Threshold
The exact point above which the agent escalates: an amount, a risk level, a confidence level. A number, not a judgement call.
Measure
What gets watched to know it works, and how often. Including how many times human review actually changed the decision.

None of the four is technical, and all four are blocking. That’s why design work comes before build work: an agent built without them demos just as well and never gets to operate.

Take the agent that’s stalled right now and answer the four questions in writing, on one page. If one of them has no answer, that’s why it hasn’t scaled. And it wasn’t the model.

Frequently asked questions

Who should own an AI agent?

The person who already answers for the outcome of the process the agent operates in, not the team that built it. If an agent decides on credit, the risk owner answers; if it decides on returns, the service owner does. The technical team answers for the agent doing what was agreed, not for what was agreed being right.

How do you decide when an agent can act without human review?

With a threshold written before it’s built: an amount, a risk level or a confidence level above which the agent escalates and below which it acts. The threshold is revised with operating data, not with a sense of how well the agent seems to be doing.

Why do most AI agents never reach production?

Rarely because of model accuracy. They stall because nobody has decided who owns the outcome when the agent is wrong. Without that decision the organisation responds by escalating every case to a person, which removes the benefit of automating it at all.

What do you measure on an agent in production?

Beyond accuracy, how often human review actually changed the agent’s decision. If that number is low and stays low, the review is costing money without adding control, and the threshold should go up.

Is legal liability the same as owning an agent?

No. Legal liability is decided by a court or the vendor contract after something has already gone wrong, and it depends on the case and the jurisdiction. The operational owner is the person the company names before the agent acts, responsible for its scope, its threshold and how it’s measured. A company can have liability fully covered in a vendor contract and still have no operational owner: those are two different questions, and both need an answer.

Let’s look at your stalled agent

From the idea to the operation

An agent reaches operation once someone defines its limits, its exceptions and who owns the outcome. That gets designed and built.

About the author

Carlos Andrés Ramírez — Transformation Director

Specialist in business transformation and reinvention. Director of Specialised Programmes and lecturer in Artificial Intelligence at UPC's Graduate School.

LinkedIn