Can an AI Vendor Train Its Model on Our Data?
It depends on what the contract says, and rarely on what the marketing page says. The useful question is not whether the vendor trains on your data, but which data it touches, for how long, and who can check. With three clauses reviewed, the answer stops being a guess.
Carlos Andrés Ramírez ·
It depends on what the contract says, and rarely on what the marketing page says. The useful question is not whether the vendor trains on your data, but which data it touches, for how long, and who can check. With three clauses reviewed, the answer stops being a guess.
Start with what usually happens. A product team tries an AI tool on real data, it works well, and a few weeks later someone in legal or security asks whether the vendor may use that information to improve its model. If the vendor also serves your competitors, the question grows: does what it learned from us end up in the product the company across the street uses?
The short answer is that a vendor can train on your data only if the contract lets it, and what is allowed varies widely between vendors and between plans from the same vendor. That is why no general answer helps. What helps is knowing which clauses to read and what to demand in each.
The question
Can an AI vendor train its model on our data?
It can if you authorise it, expressly or by omission. The second case is the one that surprises people. Many terms of use are drafted so the customer accepts broad use of what it sends, and the nuance often sits in an annex, in a linked policy, or in the difference between the paid and the free plan. A contract signed a year ago on a different plan can say something quite different from what the website says today.
There is also a second use that gets confused with training: retention. A vendor may not train on your data and still keep it for thirty days to detect abuse, or hold it longer to resolve incidents. These are separate decisions with separate risks, and they are negotiated separately.
The framework
Three clauses decide whether your data leaves your control
- 1. What counts as your data
- The contract should distinguish what you send (documents, queries, histories), what the model returns, and derived data such as usage metrics or examples marked good or bad. Training is often barred for the first and allowed for the last. If the definition of customer data is narrow, the prohibition protects little.
- 2. How long it is kept and why
- Ask in writing for the retention period and the reason. A short period for security purposes is reasonable. An open-ended one «to improve the service» is the door through which an unforeseen use walks in later. Also demand what happens when the contract ends: deletion, the deadline for it, and confirmation.
- 3. Who can check
- A promise not to train is worth what your ability to verify it is worth. Look for the right to request an independent report, the list of subprocessors that touch your data, and an obligation to warn you before terms change. Without those, the clause is a statement of intent.
The mechanism
The risk with competitors is decided by the data, not by the vendor
Imagine an insurer uploading its adjusters' notes to an AI tool to automate reports. If those notes enter general training, the model does not copy the insurer's text, but it can absorb patterns: how a loss is valued, which criterion applies to a type of claim. That edge is what a competitor could find tomorrow in the product it buys from the same vendor.
So the sensible decision is not to veto a vendor outright but to classify data before it is sent. What is public or trivial can travel on standard terms. What contains your own judgment, pricing, formulas or customer data should only travel with the no-training clause signed, the retention period agreed and the right of verification in place.
A vendor does not decide whether it learns from your data: the contract you signed does, along with your ability to check it.
BECOME
The consequence
Reviewing the contract before the pilot costs less than reviewing it after
The moment this conversation opens is usually the worst one: when the pilot already works and there is pressure to extend it. Then any awkward clause looks like a brake. Reviewed earlier, the same clause is negotiated as one more condition. What you gain is not only legal protection but the ability to tell a customer or a regulator with certainty what happens to their data.
List the AI tools your organisation already uses, including the ones a team adopted on its own. For each, find the three clauses: what data it covers, how long it keeps it, and how you can check. Where one is missing, that tool should not receive data of competitive value until it is resolved.
Frequently asked questions
Can an AI vendor train its model on our data?
Only if the contract or the terms of use you accepted allow it, which can happen expressly or by omission. The answer changes by vendor and also by plan. So read the text currently in force in your contract, not the sales material, and ask in writing for a no-training clause if the data has value.
Is a vendor keeping my data the same as training on it?
No. Keeping data is retention: it may serve security, support or audit, and it has a time limit. Training means using the data to improve the model, usually a permanent use that is hard to undo. A vendor can do the first without the second, so negotiate each under its own clause.
What if the vendor also works with my competitor?
The risk exists only if your data enters general training or a shared product improvement. With a no-training clause, short retention and a right of verification, the fact that it serves other customers weighs far less. What matters is what counts as your data and whether the contract protects it in its derived form too.
How can I check that the vendor does what it promises?
Ask for three things in writing: an independent audit report, the list of subprocessors that access your data, and a commitment to warn you before changing terms. If the vendor offers none, the promise rests on its word alone, and that should weigh in deciding what data you send.
Let's review which data can travel and under what conditions
From the idea to the operation
Scaling under control means deciding limits, oversight and traceability first. Adding them later means rebuilding.
About the author
Carlos Andrés Ramírez — Transformation Director
Specialist in business transformation and reinvention. Director of Specialised Programmes and lecturer in Artificial Intelligence at UPC's Graduate School.