LLM Security Review

AI Risk in M&A Target Due Diligence

Acquirers must audit AI systems, training data, and compliance risks that standard checklists miss.

Correspondent · · 13 min read
Cover illustration for “AI Risk in M&A Target Due Diligence”
AI Vendor Risk Analysis · September 10, 2026 · 13 min read · 2,856 words

Buying a company today means buying its AI stack, whether that company thinks of itself as an AI company or not. Models, plugins, autonomous agents, vendor integrations, the training data sitting somewhere behind all of it: these come with the deal, and most standard diligence checklists were never built to find them.

A point that gets lost in deal rooms: most companies don't self-identify as AI companies, yet they lean on AI-enabled vendor tools every single day, in the CRM, the HR system, the logistics software, the financial reporting stack. AI has spread sideways into ordinary business functions faster than legal teams have updated their questionnaires. A&O Shearman puts it plainly, too: AI issues now show up in almost every M&A deal, whether or not anyone flagged them at the start.

Speed makes this worse. Consider the UBS acquisition of Credit Suisse: UBS had under four days to assess the deal before it closed. UBS ended up setting aside close to $4 billion to cover legal and regulatory fallout afterward. That's an extreme case, forced by a banking crisis rather than ordinary deal timelines. But it's a useful baseline for what rushed diligence actually costs when it turns out wrong.

And the starting point isn't good even without AI in the picture. Sentry Tech Solutions found that only around 10% of companies run a thorough cyber due diligence assessment before acquiring a target. Layer AI-specific attack surfaces, training data provenance questions, and algorithmic bias exposure on top of a diligence process that's already thin on cybersecurity, and the gap gets a lot wider before anyone notices it.

Traditional M&A diligence runs through a familiar list: audited financials, IP ownership, litigation history, regulatory filings, employment contracts. Every category assumes a static, human-authored asset sitting still long enough to be inspected.

AI doesn't sit still. Models get retrained. Training data origins may be murky even to the people at the target company who built the thing. Outputs generated by a model may not be copyrightable at all, which matters a great deal if the acquirer is paying for that output as an asset. And vendor integrations behind a tool may have changed since the day it was deployed, quietly, without anyone updating a contract schedule.

Lexology calls this the "digital crown jewels" problem: for AI-driven companies, the real value sits in the models and the datasets, not in the underlying software code. Standard IP schedules were built to catalog code and patents, and they were not designed with AI-specific assets in mind.

Then there's the vendor opacity problem. A lot of AI capability inside a target company comes through third-party vendors, deeply wired into daily operations but rarely showing up in standard reps and warranties. The target itself may not know which specific models sit behind a vendor-supplied tool it uses every day.

Lexology's guidance on this is direct: representations and warranties need to account for the particular risks that come with developing, deploying, and using AI. That's not a matter of tweaking old software reps. It calls for language built from scratch around how these systems actually behave.

Put it together, and an acquirer relying on a standard checklist can walk away with a clean-looking report on a target whose real risk profile includes undisclosed attack surfaces, unlicensed training data, biased outputs, and open compliance violations. None of that shows up on a balance sheet.

Training data and IP ownership: where the target's AI value may rest on unlicensed ground

Plenty of AI systems get trained on data scraped from the internet or pulled from third-party sources. Squire Patton Boggs points out that copyright and trade secret protections may well apply to that underlying content, whether or not the company doing the scraping thought about it at the time.

The US Copyright Office has taken the position that AI-generated content isn't eligible for copyright protection unless it reflects what the office calls "sufficient human control." That leaves a real gap in what an acquirer can actually enforce once the deal closes, if the value being purchased is the output of a model rather than the model itself.

So the acquirer has two separate questions to answer, and they don't have the same answer just because they sound similar:

  • Does the target actually own the outputs its models produce?
  • Does the target hold proper licenses or rights to the data those models were trained on?

The contract terms worth pulling and reading line by line include licenses covering training rights, sublicensing permissions, carve-outs for text and data mining, usage restrictions, indemnities, and audit rights. Also worth checking: any indemnity the target owes to platform providers for generated outputs, pending or threatened claims tied to the model, a history of takedown requests, worst-case statutory damages exposure, and whether insurance actually covers any of it.

Hunton's review of vendor AI agreements turns up a pattern worth flagging on its own: confidentiality terms often fail to cover both the input data and the AI-generated output together, and restrictions that should stop a vendor from using a target's data to train its own models may be missing or written loosely.

If the target's flagship model turns out to have been trained on data it never had rights to, the acquirer has bought a liability wearing an asset's clothes. Standard IP schedules, built around code ownership and patent filings, were never going to catch that on their own.

Data privacy compliance embedded in AI systems and what post-acquisition regulators have shown they will do

GDPR carries no blanket exemption for scraping personal data off public websites to build a commercial AI product. It's a point a surprising number of targets haven't actually addressed, treating "the data was public" as a defense that doesn't hold up under the regulation.

The diligence review has to span several regulatory regimes at once: GDPR for anything touching EU-based data, HIPAA where medical data is involved, CPRA for California users. Each one applies both to how a model was trained and to how it processes data once it's live in production.

The EU AI Act, which took effect in 2024, adds a technical audit layer on top of all that. Among the diligence obligations the EU AI Act creates is scrutiny of how AI models used by the target were trained and whether that training meets the Act's requirements. GDPR and the AI Act intersect, and that intersection creates compliance obligations that didn't exist as a combined problem before.

The clearest warning here isn't hypothetical. The UK's Information Commissioner's Office announced its intent to fine Marriott Group tens of millions of pounds under GDPR, tied to the loss of data belonging to millions of guests. The failure traced back to IT security issues tied to Starwood Hotels systems that came under Marriott's management through the acquisition. The fine landed well after the deal had closed and the integration was complete. That's the timeline regulators actually work on.

Italy's regulator went further with a different company, temporarily suspending ChatGPT's operation in the country pending compliance fixes. The message for M&A is the same regardless of which company it happened to: regulators can shut an AI system down entirely, and that authority applies just as much to a system an acquirer inherited as one it built.

The broader pattern is worth sitting with: regulators across the EU and UK have shown they'll impose multi-million euro fines for non-compliance that happened before the acquisition, discovered after. The liability runs backward in time, not forward. Diligence has to assess the target's AI privacy posture against every jurisdiction where training data came from and every jurisdiction where outputs get delivered, not just wherever the target happens to be incorporated.

The EU AI Act's enforcement calendar and what it means for targets with high-risk AI systems

The 2024 EU AI Act is widely described as the first comprehensive regulatory framework for AI technology anywhere in the world. Penalties run up to €35 million or 7% of global annual turnover, whichever number is bigger, and market access to something like 450 million EU consumers depends on staying inside the rules.

Deal teams need to track a real calendar here, not a vague sense that regulation is "coming":

  • Article 5's prohibited-practice provisions have applied since February 2025, with full regulatory enforcement powers active from August 2025.
  • Full mandates for high-risk AI systems become enforceable December 2, 2027 for Annex III systems and August 2, 2028 for Annex I systems, per the Digital Omnibus on AI enacted July 27, 2026.
  • The European Commission proposed pushing the Annex III deadline out to December 2, 2027 as part of its Digital Omnibus package on November 19, 2025. A research note from the Cloud Security Alliance warns that companies treating this as an excuse to delay compliance spending may end up facing a badly compressed remediation window later.

High-risk categories draw the heaviest compliance burden of all: hiring processes, credit scoring, critical infrastructure management, border control, according to gdprlocal.com. Systems in these categories need third-party conformity assessment, registration in an EU database, a documented risk management system, and human oversight built in, not bolted on.

The practical diligence question follows directly from that list: does anything in the target's AI stack fall into a high-risk category? If it does, is it registered, assessed, and documented the way the Act requires? A gap here isn't a footnote. It can be a deal-stopper, or at minimum a real adjustment to price.

The US is moving in a different direction at the federal level. A January 2025 executive order, followed by a July 2025 AI Action Plan, pushed toward less federal AI regulation and fewer regulatory barriers. That gives US-domiciled targets a different posture than EU-exposed ones, though state law is filling in fast behind the federal retreat. Colorado's AI Act, repealed and replaced by SB 26-189, takes effect January 1, 2027 and regulates automated decision-making tools used for consequential decisions in employment, education, housing, lending, government services, healthcare, insurance, and legal services.

California adds a sector-specific rule worth flagging for any healthcare target: Assembly Bill 489, signed October 11, 2025, bars AI systems and chatbots from using language, titles, or design choices that suggest care or advice came from a licensed health professional when it didn't.

Defense and government-facing targets carry yet another lane entirely. Any AI system touching a government contract needs evaluation under ITAR and EAR obligations, a compliance track that runs separately from commercial AI regulation and gets missed easily by teams focused only on GDPR and the AI Act.

Minter Ellison's analysis makes a point that deserves to sit at the center of this section: AI tools trained on historical data can carry forward the biases baked into that history, and without deliberate correction, those biases compound through feedback loops rather than fading out on their own.

Here's the part that surprises people who assume AI risk needs new law to matter: existing anti-discrimination legislation already applies. Where a company's AI produces discriminatory outcomes in hiring, lending, or service delivery, the legal exposure exists right now, per Minter Ellison, with no need for AI-specific legislation to trigger it.

Two legal theories carry that exposure forward into an acquirer's hands:

  • Disparate treatment, where the AI uses or infers protected characteristics like age, race, sex, disability, or genetic information, either directly or through proxies such as geographic location or graduation year.
  • Disparate impact, where a facially neutral criterion ends up disproportionately affecting a protected group and can't be defended as job-related and consistent with business necessity.

The enforcement signal here isn't theoretical. Bloomberg Law reported that the ACLU filed EEOC charges against Aon Consulting Inc. and an employer using Aon's hiring assessments, alleging the assessments discriminated based on race and disability. The ACLU separately filed an FTC complaint against Aon over the same tools. Bloomberg Law also flagged an emerging trend worth watching closely: early signs suggest employers hit with hiring-bias lawsuits may get the chance to share liability with the vendors who built the AI tools in the first place.

That vendor liability shift matters directly for M&A. If the target's HR systems run on AI tools from a third-party vendor, its vendor contracts may carry indemnity obligations or litigation exposure nobody's disclosed yet, and all of that transfers straight to the acquirer at close.

Diligence needs to reach into the target's AI-driven decision systems across HR, lending, and customer service, and look for documented bias testing, any remediation history, and pending complaints or regulatory inquiries tied to those systems specifically.

Cybersecurity attack surfaces that AI vendor integrations introduce into the target's environment

AI systems bring cybersecurity problems that traditional security tools weren't built to catch: adversarial attacks designed to fool a model, data poisoning aimed at corrupting training data, outright model theft. Any of these can compromise not just the target company but the acquiring organization sitting behind it after close.

Prompt injection attacks, data exfiltration through plugins, agent-based connectors quietly reaching into systems they shouldn't touch: none of this shows up in generic governance frameworks or standard penetration testing. Every third-party AI integration is a possible way in, and most standard security reviews weren't scoped to look for it.

The vendor opacity problem shows up here too, in a slightly different form. A target company may not have a full, accurate map of which vendor-supplied AI models are actually running in its environment, what data those models can reach, or what permissions the agents behind them actually hold.

Sentry Tech Solutions' figure, that only about 10% of companies run a thorough cyber diligence assessment before acquiring, sets the baseline problem again here. Most acquirers walk into a deal without real visibility into ordinary cyber posture, let alone the AI-specific attack surfaces layered on top of it.

DealRoom's guide notes that AI-powered diligence tools now scan a target's digital infrastructure for vulnerabilities as part of the deal process. Worth being precise about what that covers, though: those tools generally assess how the target uses AI in its operations. They don't necessarily assess the AI system's own attack surface as a target in itself. Both audits matter, and they're not the same exercise.

A technical AI diligence review needs to cover a short, specific list:

  • A full inventory of every AI model in use: in-house builds, vendor-supplied tools, and anything embedded inside third-party software the target runs.
  • Plugin and agent permissions: what data each one can access, move, or change.
  • Connector and API security: how authentication works, what access controls exist, whether anything gets logged.
  • Evidence that deployed models have actually been tested against adversarial attacks, and what that testing found.

How to structure AI risk as its own due diligence workstream

Spread AI risk across the existing diligence tracks, and each specialist ends up seeing only their own slice of the problem. IP counsel reviews training data licenses. The cybersecurity team looks at attack surfaces. Privacy counsel checks GDPR compliance. Employment counsel reviews algorithmic bias claims. None of them is wrong to focus where they're focused, but the coordination gaps between those tracks are exactly where AI-specific risk slips through unnoticed.

A standalone AI risk workstream fixes that by assigning clear ownership across all four dimensions at once, and making sure they get assessed against each other rather than in isolation.

The workstream needs a few core pieces, mapped straight to the risk categories already covered:

  • An AI asset inventory: a complete catalog of every model, plugin, agent, and vendor integration in the target's environment. Nothing else in the review works without this as the starting point.
  • An IP and training data audit: licensing status, output ownership, indemnity terms, vendor contract language reviewed against the Hunton and Shumaker frameworks.
  • A privacy and regulatory compliance review: GDPR training-data provenance, EU AI Act high-risk classification and documentation, plus state-level obligations like Colorado's AI Act and California's AB 489 for any healthcare target.
  • A bias and discrimination audit: documented testing history, remediation records, pending complaints, and vendor contract indemnities tied to HR and lending tools specifically.
  • A cybersecurity and attack surface assessment: prompt injection exposure, plugin and agent permissions, adversarial testing evidence, controls against data exfiltration.

Reps and warranties need the same purpose-built treatment. Lexology's point holds here as much as anywhere in this piece: standard software reps don't reach model behavior, training data provenance, or output ownership. Each category above needs its own specific language in the deal documents, not an adapted version of a clause written for ordinary code.

One more thing worth sitting with before signing off on any of this: a pre-close diligence snapshot isn't the end of the job. AI vendor ecosystems keep changing after the deal closes, new model versions ship, new plugins get added, new enforcement actions land. Treating AI risk as a one-time check at close, rather than something monitored on an ongoing basis, misses the fact that the stack an acquirer bought on day one isn't the stack it's running on day two hundred.

More in AI Vendor Risk Analysis