Do AI vendors train on your data
Guide · Enterprise · 5 min read · last verified 2026-07-27
Training, in the machine-learning sense, means updating a model's parameters using examples — so "do you train on our data" asks whether the words your team pastes into an AI tool can end up shaping the model that everyone else uses. There is no general answer. It depends on the vendor, the product tier, the settings in force, and the date you ask, and the only source that binds is the vendor's current contractual documents — not a sales call, not a blog post, and not this article. What follows is a map of where the answer lives and how to read it. It is not legal advice; the documents it points to are exactly the ones your counsel should read.
Three questions hiding inside one
"Do you train on our data" bundles at least three separate questions, each governed by different language in different documents.
Training asks whether your inputs can update the weights of the vendor's models — the strongest and most commonly disclaimed use. Fine-tuning asks the narrower version: whether your data can shape smaller or auxiliary models, including classifiers used for routing, safety, or quality. Retention asks something different again: how long your prompts and outputs are stored, for what purposes, and who — including human reviewers — can read them.
The bundling matters because a vendor can truthfully say "we do not train on your data" while retaining prompts for security review, allowing staff to read flagged conversations, or using interaction data to tune auxiliary systems. That is not an accusation about any particular company; it is a fact about the sentence. The phrase underdetermines the practice, which is why the documents — not the slogan — are the unit of analysis.
Where the answer actually lives
Work the document trail in order of authority.
The data processing agreement comes first, because it is the contract that binds. If training on your inputs is excluded, the exclusion should appear here, in the DPA's own language — what a DPA is and what it governs is covered separately. A policy page can change without notice; the DPA changes by amendment.
Product terms per tier come next. The same brand name commonly spans products with different data postures: consumer and free tiers are often governed by broader terms than enterprise agreements, and a team quietly using personal accounts is operating under whichever terms it never read. Verify the tier you actually use, not the tier the vendor's enterprise page describes.
The subprocessor list tells you who else touches the data. Every name on it is a party whose own practices sit inside your risk.
The model-provider passthrough is the layer most evaluations miss. Many AI products are built on another company's models, which means your vendor's promise is bounded by its upstream provider's terms. Ask which model providers process your data and what the vendor's agreement with them says about training on that traffic. A sincere vendor promise does not cover a layer the vendor does not control.
One orientation note: this page is about your data flowing out to a vendor. The mirror-image question — where the vendor's own models and datasets came from — is a different diligence exercise, covered in data provenance requirements for enterprise AI tools. Both belong in procurement; they are answered from different documents.
What "we don't train on your data" can still leave open
Read the exclusion against this checklist of residual openings — each phrased as a question to ask, because practice varies by vendor and changes over time.
Is anything retained anyway, and for how long, for abuse or security review? Can human reviewers read your content, and under what trigger? Does the exclusion cover fine-tuning and auxiliary models, or only the flagship model? Is there a carve-out for "aggregated," "de-identified," or "usage" data — and how are those terms defined? Does feedback your team volunteers, such as ratings on outputs, travel under different terms than the prompts themselves? What notice do you get if the policy changes, and does silence constitute acceptance? What happens to commitments if the vendor is acquired?
None of these questions assumes bad faith. They exist because the sentence permits every one of them while remaining technically true — and because the only way to know a given vendor's current answer is to ask that vendor, in writing, about the tier you actually use.
Opt-out mechanics, and why defaults matter
Where training is permitted, it is often controlled by a setting — and defaults differ by product and tier, move over time, and get renamed. A toggle that exists today may sit in a different menu next quarter. The operational rule: verify the setting in the product, on the accounts your team actually uses, and get the posture confirmed in writing on the tier you contracted. A screenshot of a settings page documents a moment; it is not a contract. Where the vendor offers a written opt-out or a zero-retention option, prefer the version that lives in the agreement.
The questions to ask before you sign
Compressed into a procurement list: Is training on our inputs excluded in the DPA itself? Does the exclusion cover fine-tuning and auxiliary models? What is retained, for how long, and who can read it? Which upstream model providers process our data, under what terms? How are "aggregated" and "de-identified" defined? What notice accompanies terms changes? For early-stage vendors, fold these into the broader vetting described in the security questionnaire and the early-stage AI vendor — a young company's honest "here is what we haven't built yet" is worth more than a polished evasion.
The answer you verified is dated
Vendor data practices are a moving surface: tiers restructure, upstream providers change, policies get rewritten with notice periods measured in weeks. Whatever you verify is true as of the day you verified it, and the discipline that keeps it true is repetition — assign an owner, tie a re-check to every renewal and every terms-change notice, and record the date of each verification. The place this answer becomes operational is your own AI use policy for marketing: its approved-tools list assumes someone has done exactly this homework, per tool and per tier, recently enough to trust.