How to Audit a Decision an AI Model Made

An auditor asked why an AI model, bought from a vendor, made a particular call. What came back was a confidence score. A percentage is not a reason. It is a measure of how sure the model was that it could be wrong just the same.

Carlos Andrés Ramírez ·

An auditor asked why an AI model, bought from a vendor, made a particular call. What came back was a confidence score. A percentage is not a reason. It is a measure of how sure the model was that it could be wrong just the same.

Almost everything written about auditing AI decisions is written for whoever built the model. It explains how to read feature importance, how to produce a model card, how to keep a training log. That is real, useful information, and it serves a minority of companies: the ones running a model their own team built. Most companies are not there. They buy a platform with an interface on top and a contract underneath, and the model making the call lives on someone else's infrastructure.

That gap is not a technical footnote. It is the difference between being able to answer an auditor and not. When the model is yours, you can trace the data, the version, the criterion. When you bought it, you have exactly what the vendor agreed to hand over in the contract, and almost no AI purchase contract I have reviewed mentions the word audit.

The symptom

How do you audit a decision made by an AI model?

The answer circulating today comes from the side of whoever builds models: technical explainability, feature importance, model cards. None of those tools were built to answer what an auditor, a regulator or a lawyer actually asks, because none of them wants to know how the model thinks in general. They want to reconstruct one specific decision, made on one specific day, about one specific person or case, and know whether anyone had the chance to review it before it took effect.

  • The vendor's dashboard shows the model's overall accuracy, not the reason behind a single case.
  • The record of each decision lives inside the vendor's platform, kept for weeks, not years.
  • Nobody at the company can say which model version made a call six months ago, because the vendor updates the model without notice.
  • The purchase contract obligates the vendor to hand over the result, never the reasoning behind one particular case.
  • When a complaint or an audit lands, the only answer on hand is that the system decided that way, and that line does not satisfy a regulator.

The problem underneath

The trail an auditor asks for is not the one an AI vendor offers.

A vendor measures itself in aggregate: overall accuracy, error rate, response time. Those numbers prove the system works as a whole, and they say nothing about one specific case. An auditor does not ask whether the model gets it right almost every time. They ask why it got this one right or wrong, for this person, on this date, and whether anyone had a chance to correct it before the effect became irreversible.

And here is the part no vendor demo mentions: buy the AI instead of building it, and the trail is not yours. It belongs to the contract you signed, and if that contract did not require it in writing, it does not exist. No internal engineer can reconstruct it later, because the case data was never on your infrastructure to begin with. It sits on the vendor's, and the vendor decides how much it keeps and for how long.

Auditing an AI decision is not explaining how the model thinks in general. It is being able to reconstruct, case by case, what information it had, what it decided and who could have corrected it.

BECOME

The framework

What has to exist to audit an AI decision, bought or built?

The five pieces below get negotiated before signing with a vendor, not after an incident. Asked for afterward, no vendor hands them over for free, and one of them, the exact model version that made a past decision, cannot be reconstructed at all if nobody kept it from day one.

Case record
The exact input, the exact output and the timestamp, for each individual case. Not the model's aggregate metric: the file for that one decision.
Version
Which exact model version made that decision. A vendor that updates the model without versioning it makes it impossible to reproduce, six months later, what reasoning produced a given result.
Intervention point
The moment in the process where a person could review or reverse the decision before it took effect, and whether they did or let the model's output pass through unchecked.
Export clause
The contractual obligation, signed before purchase, to hand over the record of any case whenever the company asks, not only when the vendor chooses to share it or when the contract ends.
Internal owner
A named person, not a generic technology team, responsible for being able to reconstruct any past decision when an auditor, a regulator or an aggrieved customer asks for it.

None of the five requires building the model in house. It requires negotiating them before signing, with the same rigor applied to a service-level clause or a data ownership clause. A company that buys AI without these five is not saving the cost of building. It is deferring the cost of not being able to explain, the day someone asks, why the system decided what it decided.

Take the last important decision a purchased AI system made, a credit rejection, a candidate screened out, a fraud alert, and ask for the complete record of that one case. If all that exists is a confidence score, there is no trail. There is a black box wearing someone else's logo.

Frequently asked questions

How do you audit a decision made by an AI model?

By reconstructing the specific case, not by explaining the model in general: what data went in, which model version processed it, what came out, and whether a person had the chance to review or reverse the result before it took effect. Without those four pieces saved from the moment the case happened, the audit cannot be done afterward, no matter how well documented the model is in general.

Can you audit a decision made by an AI model the company bought rather than built?

Only if the vendor contract required it in writing before it was signed. The trail behind a purchased model depends entirely on what the vendor chooses to keep and hand over, and most AI purchase contracts never mention the word audit. Negotiating that clause after an incident is already too late, because the case data that needs reconstructing may never have been kept in the first place.

What is the difference between technical explainability and decision traceability?

Technical explainability describes how the model works in general: which variables carry the most weight, how accurate it is overall. Decision traceability reconstructs one specific case: what information the system had on a given day, what it decided and who could have corrected it. A model can have a flawless technical writeup and no individual case that can be reconstructed, and that is exactly what an auditor checks first.

Who should be able to reconstruct an AI decision when an auditor asks?

A named person with assigned responsibility, not a generic technology team and not the vendor left unsupervised. If nobody at the company can say what happened in a specific case without calling the vendor and waiting days for an answer, the company has no internal trail. It depends entirely on someone else's goodwill to answer to a regulator.

Let's design your decision trail

From the idea to the operation

Scaling under control means deciding limits, oversight and traceability first. Adding them later means rebuilding.

About the author

Carlos Andrés Ramírez — Transformation Director

Specialist in business transformation and reinvention. Director of Specialised Programmes and lecturer in Artificial Intelligence at UPC's Graduate School.

LinkedIn