AI Security Controls in Vendor Contracts and SLAs
Contracts must protect data from vendor training and map hidden AI supply chains.

Standard SaaS contracts were built for software that does one predictable thing. AI systems don't work that way, and most vendor agreements signed today still pretend otherwise. That gap is the subject of this piece: what has to change in a contract before an AI vendor gets access to real data and real workflows.
Traditional software risk has clean edges. The application does what it's configured to do. Data flows are mapped and known. The server is either up or it isn't, and an SLA that measures uptime covers most of what could go wrong. Generative AI doesn't sit still long enough for any of that to hold.
Start with training. Prompts, embeddings, and fine-tuning outputs can feed back into a vendor's model unless the contract says otherwise, and silence reads as permission. Then there's model updates: vendors push changes to weights and behavior without warning, so a system that passed acceptance testing in January might behave differently by April, with nobody re-running the test. Shared infrastructure adds another layer, since embeddings, logs, and context windows can leak across tenants if isolation isn't built in and enforced. And most AI products aren't really one vendor. They're a chain: an application layer calling a foundation model API, sitting on cloud infrastructure, sometimes routed through a separate model host. Each link is a custody point the buyer usually can't see.
Then there's the SLA problem, and it's a genuinely strange one once you sit with it. A model can report 99.9% uptime while its outputs quietly drift into bias, inaccuracy, or noncompliance. The server is up. The judgment coming out of it is wrong. Uptime doesn't measure that, and most contracts don't ask it to. The OECD AI Policy Observatory has pointed out that accountability gaps in AI deployment usually come from missing agreement on what "good performance" even means, not from clean failures anyone would notice right away.
Agentic systems push this further still. An AI agent doesn't wait for a human to click "approve." It reads data, calls tools, sends messages, and kicks off downstream workflows, often making dozens or hundreds of small decisions before anyone reviews the first one. At that point, the vendor's infrastructure has become part of the buyer's control environment, whether the contract acknowledges it or not. So what does the contract actually need to do? It needs to translate internal governance controls, the kind a security team already enforces on its own systems, into obligations a vendor is bound to meet. That's a different drafting job than a standard master services agreement, and treating it as the same one is where most of the exposure starts.
Data ownership and training rights: the clause most contracts get wrong
The single obligation worth fighting for above all others reads something like: the provider shall not use customer data to train, improve, or enhance AI models without the customer's explicit written consent. Without that sentence, the vendor's default position is permission by omission. Silence in a contract isn't neutral, it's a green light.
But "customer data" has to mean more than the files a company uploads. Scope creep happens in the derivatives:
- Embeddings generated from customer content
- Fine-tuning outputs and any resulting model weight updates
- Interaction logs and full prompt histories
Each of those can carry as much sensitive signal as the raw input, sometimes more, and a clause that only protects the original document misses all of it.
Ownership of outputs matters just as much. Who owns the insight, the recommendation, the generated report? Can the vendor take that output and use it to sharpen a model serving other customers? If the contract doesn't say, the vendor's terms of service probably already answered the question, and not in the buyer's favor.
Technical controls should sit alongside the legal clause, not replace it: logical isolation through a dedicated tenant or virtual private cloud, encryption in transit and at rest, no shared caching or embedding stores across customers. Loose data rights don't stay a competitive problem for long, either. They become a compliance one fast, since GDPR, CCPA/CPRA, HIPAA, and various sector rules all create obligations that require organizations to account for where data flows and who handles it.
Vendors will push back, often by reserving rights to train on "aggregated, anonymized" prompts. Fair enough, maybe, but the contract needs to pin down exactly what anonymization means in that vendor's pipeline, and it should require customer sign-off before any such use starts, not after.
Sub-processor transparency and the hidden custody chain
One AI vendor is rarely just one vendor. A typical stack chains an application layer on top of a foundation model API, running on cloud infrastructure, sometimes routed through a separate model host for inference. Each hop is a distinct breach surface, and each one is a party the buyer never signed anything with directly.
The IBM 2025 Cost of a Data Breach Report found that 13% of organizations reported a breach involving AI models or applications, and 97% of those organizations lacked proper AI access controls. The sub-processor chain is where a lot of that exposure lives, because access control gets weaker the further data travels from the buyer's own oversight.
A few contract requirements close most of that gap:
- The vendor keeps a current, named list of every sub-processor, along with its role and what data it can touch
- The customer gets advance notice, and a real right to object, before any new sub-processor joins the chain
- Adding a sub-processor without consent counts as a termination-for-cause trigger, not just a notice item
Sub-processor expansion without consent shows up repeatedly among the clause patterns that create lock-in years after signing, the kind buyers only discover at renewal when it's too late to negotiate from strength. That's why this belongs in the term sheet at signing, not the renewal conversation. Data residency and jurisdiction need specifying for every leg of the chain too, not just the primary vendor relationship. Practically, that means requiring the vendor to flow down equivalent security and data-handling terms to every sub-processor, and to certify that compliance on a defined schedule rather than on request.
Security baseline requirements and the certifications that actually matter
Certifications should be named in the contract, not gestured at. SOC 2 Type II, ISO 27001, FedRAMP for anything touching government work, HITRUST for healthcare, and a requirement that whatever's cited is current and active. A vendor that "used to have" SOC 2 is not the same as one that has it now.
Security measures should map to a named framework, ISO/IEC 27001 or NIST CSF, because a framework gives an auditor something to test against instead of a paragraph of marketing language.
AI systems need controls that a standard SaaS security review doesn't ask for:
- Prompt injection defenses, documented in writing, not just asserted verbally on a sales call
- Adversarial input testing, done before deployment and repeated on an ongoing basis
- Role-based access, multi-factor authentication, network segmentation
- Audit logging of AI system interactions specifically, not just application-layer events like logins and API calls
Censinet's 2025 findings put 35% of healthcare cyberattacks as stemming from third-party vendors, and healthcare is a useful stress test here precisely because the regulatory stakes are so high. But the underlying exposure isn't sector-specific. Any industry running AI vendors without a security assessment is carrying the same risk, just with different penalties attached.
Breach notification requirements, with mitigation steps included, should be set at a demanding standard in every contract, not left to statutory minimums alone. Statutory minimums vary by state and country, so treat them as the starting line, not the finish. And for AI vendors specifically, ask for disclosure of training data sources, at least by category, along with model cards laying out intended uses, known limitations, and bias testing results. Think of a model card as structured documentation of how a model behaves, covering intended uses, known limitations, and bias testing results.
SLA terms that measure AI quality, not just availability
A binary uptime SLA tells a buyer almost nothing useful about an AI system. A model can be at 99.9% availability while its outputs are biased, stale, or flatly hallucinated. Uptime and output quality are separate axes entirely, and a contract that only measures one is only half watching.
A few metrics worth adding directly into the SLA:
- AI Decision Accuracy: percentage of correct outputs against an agreed test set, with a real target attached (above 95% for something like triage decisions, for example)
Security operations offer a sharper illustration. Leading AI SOC benchmarks set 2026 targets of Mean Time to Detect under 1 hour, and organizations running SOAR-integrated AI hit MTTR figures 60% to 90% lower than teams without that automation. Worth spelling out in contract language: MTTD, MTTA, MTTR, and MTTC are not interchangeable. Some vendors report MTTA and call it MTTR, which quietly hides where the actual delay sits. "Best effort response" is not a service level. It's a way of avoiding one.
Service credits need calibrating to actual business impact, not a flat percentage of monthly fees. An AI outage that knocks out fraud detection or customer support for six hours costs a lot more than the typical credit structure covers, so the penalty should scale with what's actually at stake. As AI governance expectations evolve, the SLA is becoming more than just an operational document. It is increasingly treated as compliance evidence in its own right.
Model drift clauses and the escalating remedy structure
Model drift is quieter than an outage, and arguably more dangerous for exactly that reason. As real-world data shifts away from what a model was trained on, performance degrades gradually. The system keeps returning answers. They're just increasingly wrong.
Take an AI system built to assess mortgage applications. If it spends three months systematically undervaluing properties in certain postcodes, the service credit owed under a typical SLA might land in the low thousands of pounds. Meanwhile, the regulatory exposure and legal liability from a pattern that looks a lot like discriminatory lending could run orders of magnitude higher. SLA credits alone were never built to cover that gap, and pretending otherwise leaves the buyer holding the real cost.
A handful of contract provisions address drift directly:
- Set explicit performance baselines at signing, the benchmark the model has to hold to
- Require the vendor to actively detect deviation from expected outputs, not just track whether the service is reachable
- Mandate vendor-maintained audit logs and performance dashboards the customer can actually see
- Define retraining obligations: once performance drops below an agreed threshold, the vendor has a defined window to retrain or adjust
An escalating remedy ladder gives the relationship somewhere to go when drift shows up instead of one blunt option:
- Tier 1: the vendor notifies the customer the moment drift is detected
- Tier 2: service credits kick in if drift persists past a defined cure period
- Tier 3: human-in-the-loop review becomes mandatory once accuracy falls below a stated floor
- Tier 4: termination for cause, if the system fails a bias audit or blows through a hallucination limit for three consecutive months
None of that works if human review is left to the vendor's discretion. The contract should name which categories of decision require a person in the loop, so it isn't a judgment call made after something's already gone wrong.
Liability allocation and indemnity for AI-specific harms
Standard SaaS indemnity clauses were written for a narrower set of harms than AI actually produces. Hallucinated output used in a clinical or operational decision. IP claims tracing back to training data or generated content. Regulatory violations tied to an automated decision, touching privacy law, employment law, consumer protection. Biased or discriminatory outputs. The Mobley v. Workday litigation, where an AI vendor faces discrimination allegations under federal and state law over applicant-screening tools, shows how liability for an AI-driven decision can land on both the vendor that built the system and the customer that deployed it.
Customer-side indemnity should reach across all of that: training data claims, output infringement, confidentiality breaches, privacy violations, security incidents, regulatory violations, biased outputs, harmful content generated by the system.
Expect vendor pushback on a few fronts. Some vendors carve AI-generated output entirely out of IP indemnity. Some reserve rights to train on customer prompts regardless of what the data ownership clause says elsewhere. And standard liability caps, the kind lifted from a generic SaaS template, are frequently set well below what a copyright dispute or a regulatory enforcement action would actually cost.
The underlying principle is simple enough to state: liability should sit with whoever controls the thing that failed. The vendor controls model architecture, the training process, safety systems, security controls, and its own sub-contractor relationships. If that's where the control sits, that's where the liability belongs too. Negotiation should aim for mutual indemnities, "super caps" carved out for the categories that matter most, privacy breaches and security incidents chief among them, and explicit language preventing a vendor from capping total liability at a figure that couldn't cover a real regulatory fine if one landed.
Worth checking at contract end, not just at signing: whether data residency and model deprecation clauses are still fit for purpose given how the vendor's stack has evolved.24 research found that 62% of enterprise AI contracts lack adequate data residency and model deprecation clauses. Deprecation rights without a corresponding credit are one of several lock-in patterns that quietly compound liability exposure right when a buyer is trying to exit the relationship, not open a new front in it.
Audit rights and transparency obligations that give buyers real visibility
A standard SaaS right-to-audit clause doesn't reach far enough for an AI system. What's needed instead:
- The right to review AI compliance documentation, model cards, and bias testing results
- The right to inspect training data sources, at minimum by category and by provenance
- Access to performance dashboards and audit logs on reasonable notice, not only during a formal audit window
- Current SOC 2 Type II reports and proof that certifications are actually live, not lapsed
For anything high-stakes, hiring decisions, credit decisions, medical triage, buyers should push for third-party audits focused specifically on bias, fairness, and accuracy, with the vendor obligated to fix what's found within a defined window. A vendor that resists this is telling a buyer something worth hearing.
Transparency about change matters just as much as transparency about the current state. The vendor should notify the customer of material changes, to the model itself, to training data sources, to sub-processor relationships, before those changes go live, not in a quarterly summary after the fact. Output-IP ambiguity is one of the clause patterns that quietly locks buyers in, and audit rights reaching into IP provenance documentation are what surface that problem while it's still fixable, rather than in the middle of a dispute.
For public companies especially, SOC 2 and ISO auditors are increasingly asking for evidence that AI systems are being monitored against the standards written into the contract. The audit rights clause is what generates that evidence in the first place. Where responsibility splits between vendor and customer (training data on one side, deployment environment on another, end use somewhere else entirely) the contract should say exactly where the line falls. Leaving that ambiguous doesn't protect either party. It just guarantees the argument happens later, after something's already gone wrong.


