LLM Security Review

Third-Party Risk Management for Vendor AI Features

Enterprises miss AI risks hidden inside trusted vendors' product updates.

Reporter · · 10 min read
Cover illustration for “Third-Party Risk Management for Vendor AI Features”
AI Vendor Risk Analysis · August 29, 2026 · 10 min read · 2,233 words

A vendor bolts a copilot onto its dashboard, and nobody touches the risk file. AI features get slipped into products enterprises already trust, and the review process built to catch new vendors has no way of noticing a new feature riding in on an old one.

Why AI features don't look like risk events to TPRM teams

Standard TPRM treats risk as something that lives at the relationship level. A vendor is a company, you check that company at onboarding, and you re-check it once a year if you're disciplined about it. The whole model assumes the thing you assessed stays roughly the same shape until the next review.

AI features don't respect that assumption. They don't arrive with a contract and a security questionnaire attached. They show up in a changelog. A release note. An email from customer success announcing "exciting new capabilities." "We added AI summarization to your dashboard" gets written to read as a product upgrade rather than a risk event.

Security and TPRM teams usually aren't even on the distribution list for that kind of announcement. It lands with end users, or whoever manages the IT relationship day to day, and it just sits there. And even when someone on the risk side does catch it, the existing vendor questionnaire has nowhere to put what they found. There's no field asking a vendor to name its AI sub-features or the models running underneath them, because the questionnaire was written before any of this existed.

PwC has said plainly that traditional vendor management tools were never built to ask AI-specific questions: how the model was trained, what bias mitigation looks like, where the data actually came from. Those categories aren't skipped on purpose. They're just missing, because the template predates the problem it now needs to catch.

There's a deeper mismatch under all of this. Regular software is deterministic: same input, same output, every time, and only a code change moves that needle. AI systems are probabilistic. Output can drift or quietly get worse with no code change at all. A SOC 2 report tells you about access controls and encryption. It tells you nothing about whether a vendor's model hallucinated in a meaningful chunk of customer conversations last month, because that was never what SOC 2 was built to measure.

The risk categories that standard vendor reviews cannot see

Diagram: The Shadow AI Price Tag: What the Gap Costs. Visualizes: Visualize the financial exposure created by shadow AI using two concrete figures from IBM's 2025 breach cost data: shadow AI adds $670,000 per incident, and breaches involving shadow…

Start listing what conventional due diligence misses, and the list gets long fast.

Model drift comes first, and it's the hardest one to catch on any fixed schedule. AI systems degrade as real-world data drifts away from what they were trained on, no code change required. A model that passed its onboarding review in January can be putting out garbage by December, and an annual questionnaire just isn't fast enough to notice in time.

Hallucination is a cousin of drift but a separate problem. Generative AI gives confident, wrong answers, and in legal, financial, or healthcare workflows, the liability for acting on a wrong answer tends to land on whoever deployed the tool, not whoever built the model. Stack reputational exposure on top of that: when an AI feature says something wrong or offensive to a customer, the brand that eats the backlash is the one the customer actually knows. The model vendor two layers back stays invisible.

Bias is its own animal entirely. A biased model in hiring, lending, or healthcare can trigger regulatory attention and reputational fallout on the scale of a breach, and standard TPRM, built around security checklists and compliance boxes, has no native way to catch it.

Underneath all of this sits training data provenance, probably the least visible piece of the whole chain. What data actually trained the model, and was consent for that data properly obtained? Plenty of vendors can't fully answer that. Some won't even try. That's not always stonewalling; a lot of the time it's a structural blind spot baked into how the AI supply chain itself gets built.

Then there's the fourth-party problem. Lots of AI vendors are just wrapping someone else's foundation model, calling an external API under the hood, which means your enterprise data can pass through a model you've never reviewed and honestly can't see. Surveys of teams doing AI supply chain risk work back this up directly: external APIs and SaaS-embedded AI features rank as a top concern, just behind data sources and embeddings, cited by close to three in ten respondents. Model sourcing and provenance, the exact layer where TPRM teams spend most of their time, gets flagged by only about one in eight. So a lot of review effort is probably aimed at the wrong part of the chain.

Shadow AI is the visible symptom of everything above. Employees hook AI tools straight into production systems without asking security, because the official review process never even assessed the AI feature they actually need, so they just go around it. IBM's 2025 breach cost data puts shadow AI's added cost at $670,000 per incident, with breaches involving shadow AI averaging $4.63 million overall, above the general enterprise breach average. That gap between those two numbers is basically the price tag on a process that wasn't built to see what broke it.

What the breach and exposure data says about vendor AI risk today

Diagram: Vendor AI Breach Numbers at a Glance. Visualizes: Visualize four breach statistics that together show the scale of vendor AI risk: third-party involvement in breaches at 30% (Verizon 2025, double the prior period); 59% of breaches involve…

Vendor risk was already the single biggest category before AI showed up. Verizon's 2025 Data Breach Investigations Report put third-party involvement in breaches at 30% year over year, double the prior period. Separately, research from Wipro Cybersecurity found 59% of breaches involve an external vendor somewhere along the chain.

Add AI to that baseline and the picture sharpens fast. IBM's 2025 Cost of a Data Breach Report found 13% of organizations reported a breach involving AI models or applications. Of those, 97% lacked proper AI access controls. That statistic points to basic governance gaps nobody bothered to close, more than it points to sophisticated attackers outsmarting good defenses. Supply chain attacks accounted for close to 47% of total affected individuals in incidents during the first half of 2025, and third-party or supply chain compromise cost an average of $4.91 million per incident, well above the overall global average of $4.44 million for the year.

None of these are breaches at AI companies, either. They're breaches at ordinary enterprises where some vendor's AI feature turned out to be the unlocked door. Roughly 30% of AI security incidents trace back to supply chain compromise: infected apps, compromised APIs, third-party plugins wired into systems that were never built to expect them.

Here's something worth sitting with: most breach reporting still doesn't break out "vendor AI feature" as its own category. So these numbers are probably a floor, not a ceiling. For anyone building the case to a CFO or a board, you don't need a hypothetical worst case to make it land. The ordinary vendor relationship already sitting in your vendor register is already the main road attackers are walking down.

How regulation is shifting deployer liability onto the enterprises using vendor AI

Here's the distinction that changes the math for most companies reading this: you're probably not an AI provider. You're a deployer. You didn't build the model, but regulators are increasingly making you responsible for what happens when it runs inside your walls.

The EU AI Act is the clearest case. It entered into force in August 2024, and the date that matters most is August 2026, when deployer obligations under Article 26 kick in for high-risk uses like hiring, credit, healthcare, and public services. Article 26 asks a deployer to use the system within the vendor's stated purpose, keep human oversight from someone actually qualified to exercise it, monitor performance on an ongoing basis, retain logs for a set period, and in some cases run a Fundamental Rights Impact Assessment. None of that works if the vendor won't hand over instructions for use, technical documentation, or the logging the law requires. So vendor disclosure stops being a nice-to-have and becomes a prerequisite for staying compliant at all.

NIST's AI Risk Management Framework moved the same direction. Its March 2025 update took on generative AI risk, supply chain vulnerability, and third-party model assessment head-on, and the companion Generative AI Profile, NIST AI 600-1, spells out guidance specific to large language models. Industry guidance says roughly the same thing in different words: the old vendor model isn't enough anymore, because AI brings hallucination, drift, and supply chain complexity that need their own targeted questions, mapped to the framework's four functions of Govern, Map, Measure, and Manage. The FTC, SEC, CFPB, EEOC, and Department of Defense all reference NIST's framework now, and federal procurement is leaning toward it as the default benchmark for vendor evaluation.

ISO/IEC 42001:2023 adds another layer as an international AI Management System standard addressing AI systems that vendors and suppliers run on an enterprise's behalf. In financial services, financial-sector AI risk frameworks put real weight on third-party and fourth-party AI oversight, on top of GDPR controller duties, NIS2, and DORA's ICT risk rules.

State rules are stacking up too. State-level rules are stacking up across multiple jurisdictions, covering areas from transparency to bias audits for automated decision tools. There's no single federal framework tying all this together yet, but state and sector rules are piling up faster than most vendor contracts are getting rewritten to keep pace.

Run the thread through all of it and it lands in the same place every time: the obligation falls on the company using the AI, not just the one that built it. A vendor contract silent on AI features leaves you holding the exposure regardless of who wrote the code.

What AI-specific vendor due diligence actually needs to cover

AI-specific due diligence calls for a different questionnaire, built around the idea that each embedded AI feature is its own risk surface, with its own questions, separate from the vendor relationship as a whole.

Before you assess anything, you have to actually find it. That means building a real inventory of which vendors have added AI features since the day you first onboarded them, pulled from release notes, contract amendments, data processing addenda, sub-processor lists, and direct questions sent to the vendor. Any vendor that can't or won't name its AI features and sub-processors has already told you something. Treat that silence as elevated risk on its own, no further digging required.

From there, the questions get specific:

  • What foundation model or external API sits under this feature? Is it proprietary, open source, or licensed from a third party? What data trained it, is consent for that data documented anywhere, and does your own product data feed back into training? Can you opt out if it does?
  • How does the vendor catch drift, and what happens when performance slips? For generative features, what hallucination rates has the vendor's own testing turned up, and under what conditions did those numbers hold?
  • What human-in-the-loop controls exist, and can you actually configure them, or are you stuck with whatever default the vendor picked?
  • Has the model been tested for disparate impact across categories that matter for your use case, and can the vendor produce bias audit results that get refreshed every time the model gets retrained?

A real assessment traces the whole path: enterprise input, through whatever intermediary APIs or foundation models sit in between, to the output you actually see. Contracts should require notice the moment a sub-processor tied to an AI feature changes, not just the standard GDPR-style processor notice. And for regulatory alignment, the questions should map straight to what the law asks: can the vendor hand over the technical documentation and instructions for use that Article 26 requires of you as a deployer? Can vendor answers be mapped against NIST's Govern, Map, Measure, and Manage functions, so gaps surface systematically instead of by luck?

Why one-time assessment is structurally insufficient for AI vendor features

Deterministic software stays put until someone changes the code. AI vendor features don't play by that rule. They can shift behavior with no version bump, no contract amendment, no announcement of any kind, through retraining, through drift, through a new sub-processor quietly slotted in between one review cycle and the next.

Three failure modes live right in that gap. A vendor retrains its model and the bias or accuracy profile moves, and nobody outside the vendor notices until something breaks downstream. A new AI feature ships as an ordinary product update between annual reviews, invisible to a process that only looks once a year. A fourth-party sub-processor gets added, changing where your data actually travels, without tripping any formal amendment that would have flagged it.

Regulation is already pointing at the fix, honestly. Article 26 asks for ongoing performance monitoring and incident reporting rather than a point-in-time check, obligations a once-a-year questionnaire can't satisfy no matter how thorough it looks on the day it's filled out.

So what does continuous monitoring actually look like? Automated alerts when a vendor updates its published sub-processor list or AI documentation. Regular scans of vendor AI assets to catch new features and new data flows as they show up. Threat intelligence tuned specifically to the AI supply chain, not general vendor risk. The initial assessment still matters, but continuous monitoring is what has to run underneath it, because the thing you're assessing simply doesn't hold still long enough for a once-a-year check to matter.

Sources

  1. pwc.com

More in AI Vendor Risk Analysis