Guide / Data, Technology & Intangibles

AI Vendor Contracts: Trace Data, IP and Open-Source Dependencies

Map what enters an AI service, what the vendor may retain or reuse, what leaves it, and which software and licence dependencies remain hidden.

“You own your output” answers only one line in an AI transaction. It does not reveal what happens to prompts, uploaded files, embeddings, feedback, logs, fine-tuning data, model weights, connected systems or third-party components.

AI vendor review becomes clearer when every artifact is traced to a right, a purpose, a controller and an exit path.

Fact: third-party AI risk includes data, software and rights dependencies

The US National Institute of Standards and Technology describes the AI Risk Management Framework as voluntary. Its Core calls for organisations to map third-party software and data risks, including possible infringement of third-party intellectual-property or other rights, and to address AI supply-chain risk. NIST’s Generative AI Profile says third-party generative-AI integrations can create intellectual-property, privacy and information-security risks and points to procurement diligence, service levels and software bills of materials as possible controls.

Data-protection roles are not settled by a label alone. The UK Information Commissioner’s Office says in its AI accountability guidance that an organisation determining the purposes and means of personal-data processing can be a controller regardless of how the contract describes it. The ICO’s controller-processor guidance was under review following the Data (Use and Access) Act when checked.

An inventory is also not a licence conclusion. The US Cybersecurity and Infrastructure Security Agency describes an SBOM as a source of software-supply-chain transparency in its SBOM resources library. Compliance with open-source and other third-party licences still depends on the actual components, licences, use and distribution model.

Signal: procurement asks only whether customer data trains the model

That question matters, but a signal suggests the evidence map is narrower than the system.

  • “No training” is stated without defining retention, human review, abuse monitoring, evaluation, support or legal-compliance use.
  • Product terms, privacy notice, data-processing addendum and enterprise order contain different data-use language.
  • The vendor can change models, subprocessors or regions without meaningful notice or an exit right.
  • Output ownership is promised, while input rights, vendor materials, similar outputs and third-party claims are addressed elsewhere.
  • The business uploads customer documents without checking whether its own contract permits that processing.
  • Retrieval-augmented generation or connectors expose repositories beyond the intended dataset.
  • An open-source model name is treated as proof that every weight, dataset, dependency and deployment right is open.
  • The team cannot produce component versions, licence notices or model provenance.
  • Safety filters, evaluation or human review are not tested for the actual use case and user group.
  • The service can be suspended immediately, but prompts, evaluations, embeddings and outputs are not exportable in a usable form.

Counter-signals

The business has a current system and data-flow map; each dataset has an approved purpose and right; vendor terms are versioned; model and subprocessor changes trigger review; component and licence records are available; high-impact outputs have defined human checks; and exit has been tested. These controls create discernment, not a guarantee that an AI system is accurate, lawful or non-infringing.

Action: build an artifact-to-rights matrix

Artifact Source Permitted purpose Vendor use Retention and location Rights or licence Exit evidence
Prompts and chats
Uploaded files
Personal or confidential data
Retrieval index or embeddings
Fine-tuning or evaluation data
Model outputs
Feedback, telemetry and logs
Models and software components

Do not fill a missing cell with a sales assurance. Link it to the controlling term, technical setting, vendor response or internal approval. Mark contradictions and unassessed items.

Separate five contracts hiding inside one subscription

  1. Data use: instructions, roles, legal basis, purposes, retention, deletion, security, location, subprocessors and incident support.
  2. Confidentiality: whether prompts and outputs are protected, who may access them and how exceptions operate.
  3. IP and content: input rights, output terms, vendor tools, feedback, training, similar outputs, claims procedure and remedies.
  4. Technology supply chain: model provider, hosting, connectors, open-source and proprietary components, versions, vulnerability response and material changes.
  5. Operations: availability, performance, evaluation, human oversight, support, suspension, audit evidence, export and termination assistance.

Connect exposure allocation to the liability and insurance guide. A broad IP promise may sit under a narrow cap, exclude particular inputs or require a claims process the business cannot meet.

Review open-source evidence at component level

Request an SBOM or equivalent inventory where proportionate. Record component, version, source, licence, modifications and distribution. Ask qualified advisers which notice, attribution, source-code, patent, trademark or reciprocal obligations are activated by the actual use. “Open source,” “open weights” and “source available” should not be treated as interchangeable legal conclusions.

Test the use case, not the product category

Run realistic examples with confidential, personal, regulated and copyrighted material removed or controlled. Measure failure modes relevant to the decision: unsupported statements, leakage, harmful output, inconsistent refusal, bias, prompt injection, connector overreach and model change. Assign a human decision owner and a stop condition.

Use the EU AI Act radar and US privacy and cyber radar to identify current regulatory questions. Neither a compliance badge nor an AI policy substitutes for mapping the actual role, data and use.

Finally, simulate termination tomorrow. Can the business export prompts, outputs, configurations, evaluations and approved records; delete or return data; revoke connectors and keys; preserve required evidence; and continue the process elsewhere? Any undocumented dependency is future leverage for the vendor.

Limitations: AI rules and products change faster than the contract register

Applicable duties depend on role, jurisdiction, sector, data, users, deployment and impact. Copyright, database rights, privacy, consumer, discrimination, product safety, cyber, trade-secret and sector rules may overlap. Ownership language does not prove that output is protectable or free of third-party rights. An SBOM does not establish licence compliance or system safety.

The official sources linked above were checked on 13 August 2026. NIST states that AI RMF 1.0 is voluntary and under revision. The ICO guidance applies in the UK data-protection setting and was flagged as under review in part. Current vendor terms, model documentation and applicable laws must be rechecked at procurement, material change and renewal.

This is general information, not legal or professional advice. Law and facts vary. Consult qualified advisers for a specific situation.

Primary source

NIST AI Risk Management Framework. This source supports the identified facts; Paraveilux signals and recommendations remain interpretation.