LLM Security Review

Vendor AI Risk Assessment Template and Questionnaire Design

Standard vendor questionnaires miss the risks that actually surface when AI enters the supply chain.

Correspondent · · 11 min read
Cover illustration for “Vendor AI Risk Assessment Template and Questionnaire Design”
AI Vendor Risk Analysis · August 28, 2026 · 11 min read · 2,437 words

Most vendor questionnaires still ask about encryption, uptime, and SOC 2 attestations. Fine questions, all of them, but they miss what actually breaks when the vendor's product is an AI model: what trained it, whose data feeds it, and how its behavior shifts between versions without anyone shipping a new release. The standard template runs out of road at exactly the point where AI risk begins. This piece walks through why, and what a questionnaire built for AI actually needs to ask instead.

What is actually at stake when AI vendor assessments fall short

Vendor and supply chain compromise now costs an average of $4.91 million per breach, according to Secureframe's 2025 report. Supply chain attacks accounted for 47% of all affected individuals in the first half of 2025. Big numbers, sure, but they undersell what's happening once AI enters the picture specifically.

Shadow AI makes it worse. IBM's 2025 Cost of a Data Breach Report found organizations with heavy shadow AI use paid $670,000 more per breach on average than those with little or none. An unapproved AI tool that rode in through a vendor's product isn't a footnote in that math; it's a direct cost multiplier. Vendor involvement in breaches doubled to 30% year over year, per Verizon's 2025 report. Vendors and software supply chains stopped being an edge case a while back. They're a main road in now.

AI opens up attack surfaces a generic security questionnaire simply has no words for: prompt injection, data leaking out through model outputs, subprocessor chains that carry your data to a foundation model provider sitting three or four layers removed from the contract you actually signed. Standard vendor questionnaires rarely ask who trains on your data three steps downstream.

The legal exposure isn't theoretical either. In May 2025, the Pennsylvania Attorney General settled with Home365 over harm caused by a third-party AI tool the company had deployed. "We bought it from a vendor" didn't shield the deploying organization from liability. An incomplete AI vendor assessment isn't paperwork risk. It's event risk with a price tag, and increasingly a legal one.

Diagram: The Cost Multipliers: Why AI Vendor Risk Hits Harder. Visualizes: Visualize three compounding breach-cost facts to show how AI vendor exposure stacks up beyond a generic breach.

The regulatory environment that now requires AI-specific vendor scrutiny

Three regulatory regimes are converging on the same ask: document your AI vendor oversight, and keep documenting it.

The EU AI Act (Regulation 2024/1689) entered force in August 2024. High-risk obligations under Articles 6 through 49 were set to apply from August 2026, though a June 2026 Council amendment pushed standalone Annex III systems to December 2027. Eight categories count as high-risk: biometrics, critical infrastructure, education, employment, access to essential services, law enforcement, migration and border control, and administration of justice. Fines run up to €35 million or 7% of global annual turnover, whichever is higher. Here's the part that should reshape how procurement teams think: the organization that puts an AI system on the EU market stays legally accountable even when a third party built or runs it. Hiring a vendor doesn't move the exposure off your desk.

NIST's AI Risk Management Framework has drifted the same direction. The March 2025 update leans hard on model provenance, data integrity, and third-party model assessment. Guidance for 2025 and 2026 treats continuous monitoring of vendor APIs and open-source models as something you're expected to do, not something you check once a year and forget. The Generative AI Profile released in July 2024 (NIST AI 600-1) speaks directly to risks tied to vendor-supplied large language models. The framework's four functions, GOVERN, MAP, MEASURE, MANAGE, map cleanly onto how a TPRM program ought to be built from scratch.

In finance, DORA became enforceable in January 2025. It requires a documented risk assessment before signing a contract, another one every year after, a vendor register, contract terms covering the subcontracting chain, and annual reporting to supervisors. The U.S. Treasury's Financial Services AI Risk Management Framework, targeting February 2026, locks in independent testing, bias audits, hallucination measurement, and security testing as baseline expectations for AI vendor review at financial institutions.

What ties all four together: each one extends accountability into the vendor relationship and expects ongoing monitoring, not a form filled out once at onboarding. And there's a real gap between what regulators expect and what organizations can actually deliver: 77% are building AI governance programs right now, per the IAPP's AI Governance Profession Report, but only 1.5% say they're satisfied with how their governance function is staffed. Most of this industry is still figuring out the plane while it's already in the air.

ISO/IEC 42001 as a vendor signal and its limits as a proxy for real assessment

ISO/IEC 42001, published in December 2023, spells out requirements for an AI Management System. It's the first auditable certification built specifically for AI governance, and ISO/IEC 42006, which landed in July 2025, added an audit and certification body standard that gives the whole thing some teeth.

Big vendors are lining up for it. Amazon Bedrock, Q Business, Textract, and Transcribe got certified in November 2024. Anthropic followed in January 2025, Snowflake in June 2025, Salesforce in October 2025, ServiceNow in December 2025. Certification is moving fast from early-adopter signal toward something close to table stakes among major AI vendors. By mid-2026, "are you ISO 42001 certified or working toward it" shows up in roughly 40% of enterprise AI vendor RFPs in the EU and about 25% in North America.

But slow down for a second on what certification actually tells you. It confirms a management system exists somewhere inside the vendor's org chart. It says nothing about how a specific model behaves, what data trained it, or how exposed you are given your particular use case. Auditors are in short supply, the paperwork burden is real, and plenty of genuinely capable vendors haven't gotten certified yet just because the process eats time and money. And some vendors that do have the certification are slow to update the underlying documentation as their systems keep changing beneath it.

Treat ISO 42001 the way you'd treat ISO 27001 or a SOC 2 report: a decent pre-qualification signal, something that earns a spot in the first screening pass. It doesn't replace the domain-specific questions that come after.

The five domains a purpose-built AI vendor questionnaire must cover

Diagram: Five Domains Beyond the Standard Questionnaire. Visualizes: Show the five AI-specific assessment domains that a purpose-built vendor questionnaire must add on top of the standard TPRM baseline: (1) Model Identity & Documentation, (2)…

Layer these five on top of the standard TPRM baseline, security posture, business continuity, financial stability. They're additions. Nothing here replaces what you're already asking.

Model identity and documentation. Ask the vendor to name the exact foundation model, the version number, the fine-tuning datasets, and whether a model card exists anywhere. Find out if third-party or open-source models are sitting inside their product. This matters because 88% of organizations now use AI somewhere in the business, per McKinsey's 2025 research, and most of them are consuming AI through someone else's product rather than building their own. The model underneath stays invisible unless you make the vendor say it out loud. One red flag I'd take seriously: a vendor who can't name their foundation model, can't produce a model card, and just waves at "proprietary AI" without another word.

Training data provenance and data governance. Ask where the training data came from, whether it was licensed or consented to properly, what quality controls exist, and, this one matters most, whether your organization's data trains or fine-tunes models that later serve other customers. The risk is blunt and ugly: your data quietly improving a model that ends up serving your competitor. That's a privacy violation waiting to happen, and it can trip GDPR and the EU AI Act at the same time. Push on subprocessor depth here too. A foundation model provider sitting two or three layers behind the vendor you signed with might be touching your data and you'd never know its name.

Bias, fairness, and explainability. Ask for bias testing results broken out by demographic group, accuracy metrics by customer segment, documented human oversight on high-stakes decisions, and audit logs your team can actually pull and read. For anything customer-facing, dig into explainability standards, how disparate impact gets tested, adverse action support, complaint tracking, the actual path to human review. This isn't just good hygiene. The EU AI Act's high-risk categories, employment, essential services, law enforcement, require this documentation as a matter of law. The questionnaire itself becomes evidence in your audit trail.

Model drift and performance monitoring. Drift is what happens when a model's accuracy or behavior degrades as real-world data drifts from what it was trained on. Nothing breaks loudly; output quality erodes gradually until someone notices the numbers are off. Ask for the vendor's drift detection method, the thresholds that trigger a fix, how often they retrain, and how they notify customers when a retrained model starts behaving differently. Put this in the contract itself: does the SLA define a performance floor that accounts for drift, or does it only cover uptime? NIST's 2025-2026 guidance doesn't hedge on this: AI risk changes continuously, so continuous monitoring is the expectation, not a once-a-year checkbox.

Security and adversarial robustness. Ask about prompt injection testing, controls against data leaking through model outputs, adversarial input testing, API access controls, and an incident response plan that actually covers AI-specific failure modes. A standard security questionnaire asks about the perimeter. AI-specific questions ask what happens once the model itself becomes the attack surface. Worth sitting with: credential phishing attacks increased by 703% in the second half of 2024, driven largely by AI-generated phishing kits. The same capability vendors are selling you is getting turned right back against the supply chain that sells it.

How questionnaire structure and tiering affect what you actually learn

Not every vendor with an AI feature earns the full five-domain workup. A pre-qualification screen should sort vendors by AI exposure first. Ask two things up front: does this vendor's AI touch your data or shape a business decision, and does their use case fall under any of the EU AI Act's high-risk categories? Vendors that clear that bar move into the full assessment. Everyone else gets a lighter pass.

Evidence beats promises here, every time. For each domain, spell out exactly what the vendor has to produce: model cards, bias audit reports, penetration test results covering adversarial inputs, a subprocessor list with each layer's security posture attached. "We test for bias" isn't evidence. A third-party audit report or a published benchmark is.

Subprocessor mapping deserves its own line, honestly. Make vendors lay out the AI subprocessor chain explicitly, foundation model provider, fine-tuning infrastructure, inference hosting, and assess each layer instead of stopping at the name on your contract.

Then comes scoring. Decide up front which gaps block a deal outright, which ones get a contract with compensating controls tracked somewhere real, and which ones you're fine accepting with ongoing monitoring attached. Skip the rubric and your questionnaire results turn into a stack of answers nobody can compare across vendors, and nobody can defend later when an auditor asks why.

Build re-assessment into the design from day one; don't bolt it on later. NIST, DORA, and the EU AI Act all expect ongoing oversight. Write triggers into the contract itself: a material model update, a retraining event, a subprocessor swap. Any of those should kick off a fresh look, automatically.

Why assessment at onboarding is not enough and what continuous monitoring adds

Here's the core problem with a point-in-time assessment: AI systems don't hold still. Models get retrained. Fine-tuning datasets get swapped out. Subprocessors change hands. None of that obligates the vendor to tell you anything, unless the contract specifically says they have to.

So what does continuous monitoring cover that a one-time questionnaire, by its nature, can't reach?

Behavioral drift, for one: shifts in model output that signal something changed even though the version number on the product page never moved. New attack surfaces are another, prompt injection risks and adversarial inputs that only show up once the model lands in a new deployment context. Subprocessor changes matter just as much: a foundation model provider switching infrastructure, getting acquired, changing how it handles data, all of which flows straight through to your data whether anyone told you or not. And regulatory reclassification is its own animal; a vendor's use case can get swept into a new high-risk category under the EU AI Act, or some sector rule nobody on your team was watching for.

A governance framework on paper can't close this gap. Neither can a checklist bolted onto an existing TPRM program, no matter how well-intentioned. Closing it takes active detection of AI-specific threats, not a calendar reminder to redo the questionnaire next year. For TPRM, InfoSec, Privacy, and Legal working in sync, that's the gap purpose-built AI risk monitoring tools exist to fill: watching model behavior, subprocessor chains, and emerging attack patterns close to real time, instead of finding out six months late at the next scheduled review.

AI vendor assessment touches at least four functions, and each one carries a different piece of the risk.

TPRM owns the questionnaire itself: the lifecycle, the tiering logic that sorts vendors by AI exposure, the scoring rubric, the re-assessment cadence that keeps the whole thing from going stale. InfoSec brings the technical read, the part of the review that actually knows what a penetration test covering adversarial inputs is supposed to look like, and can spot the gap between a marketing claim and a real audit report.

Legal owns the contract language: SLA terms that tie a performance floor to drift, subprocessor notification clauses, liability allocation, the kind that matters enormously after an incident like the one Pennsylvania's Attorney General pursued against Home365. Privacy owns the data governance thread specifically, tracking whether a vendor's training practices touch personal data in ways that create GDPR exposure or break a consent boundary nobody flagged at signing.

None of these four can carry this alone. A questionnaire InfoSec writes without Legal's input misses the contract language that would make a red flag actually enforceable. A scoring rubric TPRM builds without Privacy at the table misses where the real regulatory teeth are hiding. The organizations doing this well treat AI vendor assessment as work shared across all four functions, not a form that gets routed through one team and rubber-stamped.

The checklist mindset that worked fine for static software, does it encrypt, is it up, does it have SOC 2, isn't going anywhere. It just stopped being enough somewhere along the way. AI vendors carry a kind of risk that keeps moving and reshaping itself after the ink dries, and the assessment process has to move with it or fall behind for good.

Sources

  1. atlassystems.com
  2. ampcuscyber.com

More in AI Vendor Risk Analysis