Supervise an Agent Without Redoing Its Work

"Just supervise it" is the instruction that ends with someone opening every case by hand. Nobody defined what supervising means, so everyone picks the option that leaves them covered: check everything.

Carlos Andrés Ramírez ·

An operations director put an agent in charge of matching vendor invoices against purchase orders before anything got paid. The instruction to her team was simple: supervise it. Three weeks in, every person on the team was opening the full invoice, the full purchase order and the vendor's email, exactly as they would if no agent existed at all. Nobody had told them what counted as review, so each of them picked the one option that left nobody exposed: redo it all.

The pattern isn't specific to that company. It shows up wherever someone puts an agent in charge of an outcome and asks another person to "supervise" it without saying what that word means in that particular process. When an instruction is left undefined, the safest interpretation wins, and the safest one is almost always the most expensive: review the work as if the agent had never touched it.

The time that was supposed to be saved doesn't disappear. It moves. It shifts from the column labelled execution, which someone was tracking, to the column labelled review, which nobody counts at all. A company can spend a full year believing it automated something while the same number of hours, filed under a different name, keeps coming out of the same payroll.

The symptom

Why does reviewing an agent end up costing as much as doing the work?

It shows up before anyone says it out loud. The reviewer takes nearly as long approving a case as they would solving it from scratch. They chase the source figure even though the agent already cited it. They re-add the total even though the agent already added it. And if a case slips through half checked, they don't sleep well, because nobody has told them what happens if that one turns out to be the case that went wrong.

  • The reviewer opens the full case every time, not only when the agent flags that something doesn't add up.
  • Nobody has written down what "review" means in that process: reading the conclusion, or rebuilding the calculation from the underlying data.
  • The agent never states how confident it is in its own answer, so the obvious case and the doubtful one get the same amount of attention.
  • The only figure anyone tracks is how many cases the agent processed, never how many the reviewer actually corrected.
  • The hours saved in execution and the hours spent on review never land in the same ledger, so the project always looks like it paid off.

The problem underneath

Reviewing everything is redoing the work under a different title.

Checking everything an agent does requires understanding each case as thoroughly as if it had been done by hand. That isn't supervision, it's execution with an extra step bolted on. And the alternative, checking a portion and trusting the rest, is frightening for a specific reason: nobody has said who answers for the case that wasn't checked if that turns out to be the one that failed. Until that question has an answer, reviewing everything is the only defensible position, even if it costs back every hour the agent was supposed to return.

Supervising isn't reading a result until you feel as certain as if you'd done it yourself. It's deciding, in advance, how much doubt you're willing to accept without looking.

BECOME

What changes

Four decisions that separate supervising from redoing.

Sample, not census
Review a portion chosen by method, random and stratified across case types, not every case that comes in. What isn't checked one by one isn't left uncontrolled: it sits inside the sample that was measured.
Flagged exceptions, not routine cases
The agent states its own uncertainty, and that's where a human eye belongs. A case the agent resolves confidently, in line with the expected pattern, doesn't need the same attention as one that falls outside it.
A trail, not a repeat
The agent writes down which figure it checked and which rule it applied to reach its answer, and the reviewer audits that trail instead of recalculating the result from raw data. Auditing a decision is quick. Redoing it isn't.
A trend, not a snapshot
Track the correction rate on the sample over several weeks, not whether today's case turned out fine. An agent earns trust through sustained behaviour, not case by case.

None of the four calls for a better model. They call for someone with authority over the process to decide how much doubt they're willing to accept unseen, put it in writing, and stop asking their team to guess where the line sits.

Ask whoever supervises an agent today how many of the cases that reach them get opened in full. If the answer is "all of them", there's no supervision happening: the same work is being done twice, and only one of the two times shows up on a payslip.

Frequently asked questions

How do you supervise an AI agent without redoing the work yourself?

By reviewing a sample chosen by method, not every case it produces. The sample weights toward whatever the agent flags as less certain and fills in with a random pick from the rest, so nothing the agent believes is settled but isn't gets missed.

How many of an agent's decisions actually need reviewing?

There's no fixed number that works across every process: it depends on how much damage a mistake causes and how often cases fall outside the expected pattern. What stays constant is the method: a random sample to watch the whole, and a full check on everything the agent flags as uncertain.

What does it mean for an agent to leave a reasoning trail?

It means the agent writes down which data it used and which rule it applied to reach its answer, not just the final result. Auditing that trail takes minutes, because it checks a chain of steps already written out. Redoing the work from raw data, with no trail to check against, takes as long as doing it the first time.

How do you know whether an agent's supervision is actually working?

By measuring, on the reviewed sample, how often a person genuinely corrects something the agent decided, and tracking that figure over several weeks rather than judging a single case. If the correction rate stays low and steady, the review threshold can go up. If it climbs, the threshold needs to come down before the mistake reaches a customer.

Let's design your agent's supervision

From the idea to the operation

An agent reaches operation once someone defines its limits, its exceptions and who owns the outcome. That gets designed and built.

About the author

Carlos Andrés Ramírez — Transformation Director

Specialist in business transformation and reinvention. Director of Specialised Programmes and lecturer in Artificial Intelligence at UPC's Graduate School.

LinkedIn