ATLAS/BRIEFINGLaw, organized for consequential decisions.

PRIV-04 Privacy, Cyber & AI Data Under Duty Federal + state overlay

Contracting With AI Vendors: Training Data, Output Rights, Security, and Liability

Buying an AI system transfers your data and imports someone else's legal exposure. This brief works the seven terms that decide who carries that risk, with realistic fallback positions.

Technical diagram marking this brief's subject

Briefing in 60 seconds

  1. Default vendor terms often permit training on customer inputs; the restriction must be written, cover outputs, and bind subprocessors.
  2. Output ownership is assigned by contract, but assignment cannot create copyright the law does not grant to purely machine-generated material.
  3. IP indemnity is the term most negotiated and most conditioned; read the exclusions, caps, and required-use conditions before relying on it.
  4. Model change management and evaluation rights matter because an unannounced model update can silently break a validated workflow.

Controlling variables

Facts
What data actually leaves your environment — prompts, documents, retrieval context, telemetry, and human feedback — since each category is treated differently.
Contract terms
Whether training restrictions, retention limits, and indemnity sit in the master agreement or in an acceptable-use policy the vendor can amend unilaterally.
Status
Whether the deployment is internal-only, customer-facing, or embedded in a regulated decision, which changes the required controls and disclosures.
Jurisdiction
Which state privacy statutes and sector regulators reach the use case, since processing terms and assessment duties differ across them.
Timing
Whether the vendor can change models, defaults, or subprocessors mid-term, and how much notice you receive before the change takes effect.

General legal information about United States law. Not legal advice, not representation, and no attorney–client relationship is created by reading it. Rules differ by jurisdiction and change — verify against the official sources listed below.

An AI purchase is a data transfer with a software invoice attached. The commercial terms rarely decide the outcome; seven contract mechanics do. Whether your inputs can train the vendor's models. Who owns and can use the outputs. Whether the vendor's IP indemnity actually covers you. How security and subprocessors are controlled. How long prompts persist. What happens when the model changes. And whether you may test the system yourself.

None of this is exotic drafting. It is ordinary technology contracting applied to a product whose behavior changes over time and whose training history you cannot inspect. The negotiation succeeds when you start from what data leaves your environment rather than from the vendor's paper.

Map the data flow before the redlines

Ask the deployment team four questions and write the answers down. What is sent to the vendor — prompts alone, or attached documents, database records, and retrieved context? What comes back, and where is it stored? What telemetry is collected, including logs, latency data, and flagged conversations? And what human feedback is captured, since thumbs-up ratings and correction text are training signal with a friendly interface.

That map determines which terms you actually need. A summarization tool operating on public marketing copy needs light controls. A support assistant that reads customer records carries privacy obligations, and if any of those records involve biometric templates or minors, the regimes covered in biometric privacy laws and children's online privacy apply on their own terms regardless of what the AI contract says.

Training data and improvement rights

The single term customers most often assume and least often verify is whether the vendor may train on their content. Enterprise offerings frequently disclaim training by default while consumer or self-service tiers permit it, and the same vendor may do both. The distinction usually lives in an online policy rather than the signed agreement, which is precisely the problem.

A workable training-data clause does four things. It prohibits use of customer inputs and outputs to train, fine-tune, or otherwise improve models made available to anyone else. It defines "customer data" to include prompts, attachments, retrieval context, outputs, and human feedback, because a narrow definition covering only "Customer Content" leaves feedback and telemetry outside. It binds subprocessors and any upstream model provider on the same terms. And it lives in the executed agreement, with any conflicting online policy expressly subordinated.

Expect two vendor counter-positions. The first is a carve-out for aggregated or de-identified data used for service improvement; that is often acceptable if de-identification is defined and re-identification is prohibited. The second is abuse monitoring, where the vendor retains prompts to detect misuse. That is a legitimate need, and the fallback is a bounded retention window, access limited to a small trained team, no human review except on defined triggers, and no use of monitored content for training.

Negotiation register: vendor default, target position, and realistic fallback
TermCommon vendor defaultTarget positionFallback that usually closes
Training on inputs Permitted, or disclaimed only in an amendable policy Prohibited in the agreement for all customer data, binding subprocessors Prohibited except aggregated de-identified telemetry, with re-identification barred
Output rights Customer owns "to the extent permitted by law," with no non-assertion Assignment of vendor rights plus a covenant not to assert against customer use Assignment plus acknowledgment that similar outputs may be produced for others
IP indemnity Capped at fees paid, excluding outputs Uncapped defense and settlement for third-party IP claims arising from the service and its outputs Supercap on IP claims, conditioned on using current models and stated safety features
Retention of prompts Indefinite, or "as needed for the service" Zero retention beyond the session, with deletion certified Fixed abuse-monitoring window of days, then automatic deletion
Model changes Vendor may modify the service at any time Advance notice of model or default changes, with a pinned version available Notice plus a deprecation window and a documented rollback path
Evaluation and audit Benchmarking prohibited Right to test, red-team, and publish internal results Testing permitted, external publication subject to notice and factual accuracy

Who owns the output, and what that is worth

Contracts can allocate whatever rights the parties hold, and most vendors will assign their interest in outputs to the customer. What a contract cannot do is manufacture intellectual property that the law does not recognize. The U.S. Copyright Office has taken the position that copyright protects human authorship and that material generated without sufficient human control is not registrable, while human-authored contributions to a work involving AI assistance may be. The Patent Office has taken a parallel position on inventorship, requiring a significant human contribution and rejecting an AI system as a named inventor.

The practical consequences are concrete. Marketing copy generated with minimal human input may be unprotectable against copying by competitors, which matters if the asset is meant to be exclusive. Code produced by an assistant is protectable to the extent of human authorship, and the record of that authorship is worth keeping. And an assignment clause with no non-assertion covenant leaves the vendor free to assert its own rights later, which is why the covenant matters more than the assignment.

Ownership language should also acknowledge reality: the same model will produce similar outputs for other customers. If a term promises exclusivity in outputs, it is either meaningless or it is a promise the vendor cannot keep. Where genuine exclusivity is required — a licensed voice, a proprietary style, a trained-on-your-corpus model — that belongs in a separate grant drafted with the discipline described in IP license agreements, including whether the grant is an exclusive license or something weaker.

Indemnity that survives contact with a claim

Major vendors now offer IP indemnities covering third-party copyright claims arising from outputs. Read them as conditional promises rather than coverage. The conditions typically include using current model versions, not disabling built-in filters, not providing infringing material as input, and not deliberately prompting for a protected work. Each condition is a factual defense the vendor can raise when a claim arrives.

  • Indemnity excludes the thing you use. Coverage may extend to a flagship model but not to fine-tuned variants, open-weight deployments, or third-party plugins. Confirm the covered service list matches your actual configuration.
  • Cap swallows the promise. An indemnity limited to twelve months of fees is not risk transfer for a claim seeking statutory damages. Push IP claims outside the general cap or into a supercap.
  • Control of defense is unclear. Specify who selects counsel, who may settle, and whether the vendor can settle on terms requiring you to stop using the service.
  • Confidentiality does not cover prompts. Many agreements protect "Confidential Information" but define it around disclosed business terms. State expressly that prompts, attachments, and outputs are confidential and are not vendor data.
  • Security terms are generic. A cloud-era security exhibit rarely addresses prompt logs, vector stores, or retrieval indexes, all of which can contain the same regulated data as the source system.
  • Subprocessors change silently. If the vendor routes to an upstream model provider, that provider's terms govern part of your data path. Require a maintained list, advance notice of additions, and flow-down of your restrictions.
  • Incident duties do not match your own. Notification windows should be short enough to let you meet your regulatory clocks, and the vendor should cooperate with a privileged investigation rather than publishing findings unilaterally.

Standard commercial indemnification for confidentiality breaches and data-protection violations should sit alongside the IP indemnity, not inside it. The two cover different failures and are frequently capped differently.

Retention, deletion, and model change control

Retention is where privacy commitments quietly fail. Prompts persist in logs; retrieval systems copy source documents into embeddings; evaluation sets accumulate real customer text. Contract for deletion at each layer, require certification on termination, and ask specifically what happens to embeddings derived from deleted documents, since a vector index is a derivative copy that deletion scripts often miss.

Model change control is the term teams regret omitting. Vendors deprecate versions and adjust defaults, and a system validated against one model can behave differently the next quarter without any change on your side. Negotiate advance notice, a pinned version for a defined period, and a deprecation window long enough to re-test. Pair that with evaluation rights: the ability to run your own accuracy, bias, and safety testing, and to use the results internally without a confidentiality clause that prevents you from telling your own regulator what you found.

Change discipline: treat a model version change like a production dependency upgrade. Re-run the evaluation suite before the new default takes effect, and keep the results, because they are the evidence that the system was monitored.

For governance structure around all of this, the NIST AI Risk Management Framework is the reference most U.S. organizations use. It is voluntary and not a legal standard, but its govern, map, measure, and manage functions give a defensible shape to documentation, and its generative-AI profile addresses risks specific to these systems. Mapping contract terms to framework functions makes gaps visible: a measurement obligation with no evaluation right in the contract is a control you cannot actually perform.

Questions the desk gets

The vendor says it does not train on our data. Is that enough?

Only if it is in the signed agreement, defined broadly enough to include attachments, outputs, feedback, and telemetry, and flowed down to subprocessors and upstream model providers. A statement in a marketing page or an amendable policy is a representation the vendor can change. Ask for the term in the contract and ask what the abuse-monitoring exception retains.

Can we own the copyright in what the model produces?

You can own whatever rights exist and whatever the vendor assigns. Whether copyright exists at all depends on human authorship, and the Copyright Office's position is that purely machine-generated material is not protectable while human-authored elements can be. For assets that must be exclusive, plan for meaningful human authorship and keep a record of it rather than relying on a contract clause.

Do we need a separate agreement for a model that runs in our own environment?

The data-flow risk drops, but licensing risk rises. Self-hosted and open-weight models come with license terms that may restrict field of use, require attribution, or impose downstream conditions on derivative models. Read the model license as you would any software license, and confirm whether the vendor's IP indemnity extends to that deployment mode. It frequently does not.

How should we handle a vendor that will not negotiate its standard terms?

Change the deployment instead of the paper. Restrict what data reaches the service, redact identifiers before transmission, disable feedback capture, and confine the tool to use cases where the unmitigated terms are tolerable. Document the decision and the compensating controls, because an accepted risk that was analyzed reads very differently from one that was ignored.

Sequencing the negotiation

Work in this order: data-flow map, training restriction, retention and deletion, security and subprocessors, confidentiality of prompts, output rights and non-assertion, indemnity, then change control and evaluation rights. The early items reduce the amount of exposure the later items have to cover, and a strong training restriction makes several downstream fights unnecessary.

Before signing, confirm three things independently of the vendor's summary: which entity actually processes the data, what the acceptable-use policy says today and whether it can change, and whether the indemnity's covered-service list names the configuration you intend to run. Where the deal also allocates commercial risk through representations and remedies, the analysis in representations and warranties is a useful companion, and further work across this area sits on the Privacy, Cyber & AI desk. If a security event later reaches the vendor's environment, the first-hours posture in data-breach response governs.

Sources

  1. NIST — AI Risk Management Framework
  2. U.S. Copyright Office — copyright and artificial intelligence materials
  3. U.S. Patent and Trademark Office — inventorship and AI-assisted inventions guidance
  4. Federal Trade Commission — Privacy and Security business guidance

Atlas Research Desk

ATLAS briefs are researched and edited by the Research Desk, an editorial organization — not attorneys acting for you. Method and limits: editorial method · source standards · corrections.