LLM Security Review

Enterprise AI Security Platforms for LLM Prompt Injection and Data Exfiltration Protection

Legacy security tools can't see attacks hidden inside LLM context windows and agent chains.

Reporter · · 15 min read
Cover illustration for “Enterprise AI Security Platforms for LLM Prompt Injection and Data Exfiltration Protection”
Data Exfiltration · September 30, 2026 · 15 min read · 3,294 words

Enterprise LLM deployments are getting attacked through a hole that traditional security tools were never built to see: the context window itself. Prompt injection and the data exfiltration it enables require purpose-built platforms that cover runtime validation, indirect injection through retrieval, and agentic tool-calling all at once, because traditional security tooling was never designed for this attack model.

Why the enterprise LLM attack surface is different from what traditional security tools were built for

Start with the architecture. An LLM takes in a system prompt, a user's message, and whatever content it retrieves from documents, emails, or web pages, and it processes all of it as one undifferentiated block of text. There's no wall between "instructions the developer trusts" and "data a stranger wrote." That's the whole problem in one sentence, and it's a design flaw baked in from the start. It's how the model works.

Security folks who've been around a while will recognize the shape of this. It rhymes with SQL injection, where code and data got mixed in a database query and attackers exploited the seam. But the scale here is different, and so is the surface. SQL injection had a bounded target: a query, a database, a known set of inputs. An LLM's "input" can be a calendar invite, a Slack message with hidden instructions that can trick an AI assistant into leaking data, or a document uploaded to a trusted collaboration tool. The seam is everywhere.

TrueFoundry's 2026 buyer's guide lays out three assumptions baked into traditional security tooling that simply don't hold for agentic systems. First, the assumption that a human initiates a discrete, single action. Agents break that by chaining multiple steps together with no person in the loop watching each one. Second, the assumption that an action has a clear start and a clear end that a reviewer can catch mid-flight. Agentic chains can run an entire exfiltration sequence in seconds, long before any human notices anything is wrong. Third, and maybe most overlooked: the assumption that identity maps to a person. Agent credentials, OAuth tokens, API keys, database credentials, don't correspond to any individual employee sitting in an IAM system. Who do you revoke access from when the "user" is a process?

This is why specific tool classes fail in specific, almost embarrassing ways. A web application firewall inspects HTTP request signatures. It has no concept of an agent's reasoning chain, so it simply can't see the attack happening. Data loss prevention tools scan for patterns in outbound files and messages, but they don't inspect the content sitting inside an LLM's context window. Cloud access security brokers attribute actions to identities, but when a tool call originates from an autonomous agent process rather than a logged-in user, there's no identity to attribute it to.

The OWASP Top 10 for Agentic Applications, published December 2025, ranks agent goal hijack, tool misuse and exploitation, and identity and privilege abuse as the top three risks facing agentic systems, and none of them have adequate coverage in any of the traditional tools just described Cloud Security Alliance. That's not a gap at the margins.

And this isn't theoretical anymore. Roughly 78% of organizations now use AI in at least one business function, which means the exposed surface is already running in production, not sitting in a lab somewhere waiting to be tested BlackFog. Cisco's State of AI Security report found that 83% of organizations plan to deploy agentic AI, but only 29% feel ready to do it securely. That gap between ambition and readiness is widening. It's widening as deployment speed keeps outpacing the security tooling meant to catch up with it.

The five-layer attack surface enterprise AI agents expose

Enterprises evaluating vendors keep making the same mistake: treating "AI security" as one undifferentiated category BlackFog. It isn't. TrueFoundry's buyer's guide breaks the surface into layers, each with its own threat types and its own specific gap in legacy coverage.

Layer one is identity: agent credentials, OAuth tokens, API keys, session identity. The threats here are credential theft, privilege escalation, and lateral movement, and traditional privileged access management or identity governance tools only partially cover non-human agent identities. Layer two is the model and prompt layer itself, covering LLM calls, context window content, and how prompts get constructed. Threats include prompt injection, jailbreaks, sensitive data leaking into prompts, and manipulated model output, and no traditional tool inspects LLM input or output at all.

Layer three is tools and MCP: MCP server connections, tool invocations, and external API calls; threats include tool poisoning, supply chain compromise, unauthorized tool access, and exfiltration via the tool chain, and traditional tools have no visibility into MCP invocations. Layer four is runtime behavior, covering multi-step task execution and state that persists across sessions. The threats are anomalous action chains, memory poisoning, and attacks that propagate from one agent to another, and there's no behavioral baseline for agent action sequences in any legacy tooling. Layer five is compliance and audit: the trail, the data governance record, the regulatory evidence. Threats include gaps in the audit trail, GDPR and HIPAA data flows going untracked, and missing attribution, and existing compliance tooling only covers this partially.

Why does this five-layer breakdown matter practically? Because a single injection doesn't stay contained to one action. The context window carries state across every tool invocation in a session, so one successful injection early on can contaminate everything that follows. A security team evaluating vendors needs to map each platform's coverage against these five layers explicitly, not just take "we cover AI security" as an answer.

Prompt injection: the mechanics, the scale, and the reasons it keeps succeeding

Prompt injection has held the number one spot on the OWASP Top 10 for LLM Applications for two editions running Cloud Security Alliance. That consistency should tell you something: this is a well-documented pattern researchers have mapped in detail. It's a problem that has scaled straight into production and stayed there.

The numbers on how often it works are uncomfortable. Attack success rates run 50 to 84% depending on the system's configuration and how many attempts an attacker gets, and in agentic systems that auto-execute actions without human confirmation, success rates hit that 84% ceiling Vectra AI. Meanwhile, only 34.7% of organizations have deployed dedicated defenses against prompt injection specifically, according to OWASP. Do the math on that gap. The majority of production LLM deployments right now have no dedicated defense against the single most successful attack vector against them.

Split the attack into its two forms, because they behave very differently. Direct injection is when an attacker types override instructions straight into their own input field, something like "ignore previous instructions." It's the easier one to catch: detection rates exceed 70% in filtered environments SQ Magazine. Indirect injection is the harder problem. Instructions get hidden inside a document, an email, a webpage, a calendar invite, or a database record that the LLM later retrieves on its own, with no attacker typing anything directly to the system Vectra AI SQ Magazine. About 62% of successful exploits used this indirect pathway, and over half of them slipped past standard prompt filtering entirely Vectra AI SQ Magazine. Web-based indirect injection alone accounts for nearly 40% of all LLM security incidents SQ Magazine.

That's a sobering line coming from a national security agency. SQL injection got largely solved through parameterized queries, a clean architectural fix. There may not be an equivalent clean fix here, because the fix would require the model to stop treating instructions and data as the same thing, which is close to asking it to stop being a language model.

Recent research out of arXiv (paper 2601.09625) reframes prompt injection using the language of a proper malware kill chain: seven stages running from initial access through privilege escalation, reconnaissance, persistence, command and control, lateral movement, and finally action on objective. And the maturity of these chains is climbing fast. Attacks that reached four or more stages of that kill chain went from zero in 2023, to 7 in 2024, to 15 across 2025 and 2026 The Promptware Kill Chain: How Prompt Injections Gradually Evolved In…. That's a steady, deliberate escalation. That's attackers getting organized.

The agentic layer amplifies all of it. OWASP's LLM Security Report clocked a 340% year-over-year surge in prompt injection attacks overall, with multi-hop indirect attacks specifically up 70% year-over-year. And this isn't scattered, opportunistic poking around. Forcepoint identified 10 distinct indirect injection payloads spread across separate domains, and Unit 42 mapped 22 separate payload-delivery techniques currently active in the wild. That's tooling. That's infrastructure being built and reused. UK NCSC has warned this class "may never be totally mitigated in the way that SQL injection attacks can be". Agentic amplification is stark: 40% of AI agent frameworks contain exploitable prompt injection flaws in tool-execution logic, and autonomous agents calling APIs exhibit up to 2.5x higher risk exposure than standalone models.

Diagram: Promptware Kill Chains Are Escalating Fast. Visualizes: Show the year-over-year growth in multi-stage promptware kill chain attacks reaching four or more stages: zero such attacks in 2023, 7 in 2024, and 15 across 2025–2026.

How prompt injection becomes data exfiltration: the linked threat chain

Injection and exfiltration aren't two separate incidents that happen to occur near each other, which matters most for anyone thinking about business risk rather than just technical curiosity. Injection is the entry point, and exfiltration is the most common "action on objective" documented across promptware kill chains. The attack keeps going once the injection succeeds, just getting started. It's just getting started.

Once an injection lands, attackers tend to go after one of two things: hijacking the agent's instructions, or pulling data out. Data exfiltration carries the heavier consequence of the two, because once information leaves the building, there's no undoing it. What actually gets exfiltrated is a broad list: system prompts, training data, API keys, proprietary documents, patient records, compliance records, essentially anything reachable either in the context window or through a tool call the agent has access to. If an LLM leaks patient data or trade secrets, that's not just an embarrassing headline. It can trigger HIPAA or GDPR violations, litigation, and a real loss of user trust that doesn't come back quickly.

One especially unsettling thread of research, presented at ACL, looked at backdoored tool-use agents specifically. Across three separate model families and three different domains, trigger activation for backdoor-driven exfiltration exceeded 94%. And the triggers didn't need to be elaborate. Simple two- or three-word conjunction phrases produced mean activation rates above 96%, with false-positive rates under 0.3%. A two-word phrase reliably triggers data theft over 96% of the time, almost never firing by accident. That's a system already operating at scale in production. That's a nearly silent, nearly perfect attack mechanism.

But what about the guardrails everyone's already deploying at the retrieval stage? Research using a reranker-aware response rewriter got past retrieval-stage defenses built on LLM Guard or NeMo Guardrails 81.2% to 86.7% of the time. Guardrail-only approaches, in other words, are not sufficient on their own. This points to an architectural truth that keeps surfacing throughout this whole space: a defense sitting only at the prompt input layer will miss exfiltration happening through tool calls, through retrieved context, or through a cleverly rewritten response. Coverage needs to span input, retrieval, output, and tool execution together, or it isn't really coverage.

What confirmed production incidents reveal about the actual attack patterns

Theory is one thing. What's actually happened in production tells a sharper story.

EchoLeak, tracked as CVE-2025-32711 with a CVSS score of 9.3 and disclosed in June 2025, hit Microsoft 365 Copilot Securance. An attacker sends a crafted email containing hidden instructions. The recipient later asks Copilot to summarize their inbox, and the AI silently exfiltrates sensitive documents to an external server, all without a single click from the victim and without any malware involved Securance. Zero-click. That phrase alone should get every security team's attention.

Slack AI had its own indirect injection incident: hidden instructions embedded in an ordinary-looking Slack message trick the AI assistant into inserting a malicious link, and clicking it sends data out of a private channel to an attacker's server. The attack surface here is just... a normal message, in a tool employees trust every day.

Bloomberg reported that a hacker used Anthropic's Claude to steal sensitive data belonging to the Mexican government. That incident matters for one reason above all others: it confirms this threat isn't confined to sloppy or misconfigured enterprise deployments. It reaches sophisticated, well-resourced targets too.

ServiceNow disclosed its own vulnerability, nicknamed "BodySnatcher" and tracked as CVE-2025-12420, in its Now Assist AI agent, patched on October 30, 2025. And scale matters here too: GitHub Copilot has reached adoption across 90% of Fortune 100 companies, and one Copilot prompt injection flaw was rated CVSS 9.6 Cloud Security Alliance. When a tool is that deeply embedded across that much of the corporate world, a single flaw doesn't stay small. It scales with the deployment.

Of 8 major AI-related incidents documented from January through April 11, 2026, only one received a CVE identifier: CVE-2025-59528, tied to Flowise. Incidents in the round-up included GrafanaGhost's indirect-injection exfiltration, an identity abuse case in Vertex AI dubbed Double Agent, and a remote code execution flaw in Flowise. Unit 42 reported what it called the first documented large-scale indirect prompt injection attacks in the wild, including techniques for evading ad review and leaking system prompts on live commercial platforms. CrowdStrike's 2026 threat reporting documented prompt injection attacks against more than 90 separate organizations.

Not every exfiltration incident involves a hacker at all, though. Some of it is just employees being careless. Engineers at Samsung reportedly pasted confidential source code into ChatGPT, and in response, both JPMorgan and Goldman Sachs restricted employee use of ChatGPT, citing standard third-party software compliance concerns about staff sharing sensitive information. That insider path doesn't get stopped by a perimeter defense. It needs governance controls aimed at how people use these tools day to day.

Pull back and look at the pattern across every one of these incidents: most of them don't fit the tidy CVE model at all. They come from misconfiguration, excessive agency, and indirect injection, things that are invisible to signature-based detection. If a security team is waiting for a CVE alert to tell them something's wrong, they're going to be waiting for something that, in most cases, never arrives.

What a purpose-built platform must cover across the full threat surface

Given all of that, what does an actual, working platform need to do? Synthesizing evaluation criteria from Maxim AI's Bifrost framework, TrueFoundry's buyer's guide, and OWASP's coverage requirements points to five capability dimensions that matter.

Runtime input and output validation comes first: catching prompt injection, jailbreaks, indirect injection buried in retrieved content, and exfiltration attempts riding out in the response, on both sides of every single inference call. Access control and governance is second, restricting which agents, teams, and applications can reach which models and tools, with spend controls and rate limits enforced at the moment of execution rather than only at the provisioning stage. Third is model supply chain security: scanning model artifacts themselves for backdoors, malicious code, or deserialization exploits before anything goes into production, not relying on runtime protection alone to catch it later.

Fourth is agentic tool governance, controlling exactly which MCP tools an agent is allowed to invoke, enforced both at inference time and again at the moment of tool execution, because schema-level checks and execution-level checks catch different things. Fifth is the audit trail: immutable records of every inference, every tool call, every guardrail evaluation, and every policy decision made, which regulated industries need for SOC 2, GDPR, HIPAA, and EU AI Act documentation.

None of these five hold up alone. Recall that ACL 2026 finding, guardrail-only retrieval defenses got bypassed 81.2% to 86.7% of the time. That's the strongest argument available for layering detection rather than betting everything on one control point. Deployment model matters too, and it's not a minor checkbox. SaaS, self-hosted, in-VPC, and air-gapped options each carry different implications for data residency and audit access, and regulated industries typically need self-hosted or VPC-isolated setups to satisfy their compliance obligations.

Latency is the unglamorous constraint that decides whether any of this survives contact with real users. Runtime classifiers add processing overhead to every single request that passes through them, and sub-100ms overhead paired with high precision is roughly the threshold for production viability. Push past that, or let false positives pile up, and users start finding workarounds, which quietly defeats the whole point of having the defense.

Vendor AI is another blind spot, especially now that 78% of organizations use AI in at least one business function, so the exposed surface is already production-scale, not theoretical BlackFog. Enterprises are running a mix of homegrown and third-party LLMs. They're deploying models, plugins, agents, and connectors built by outside vendors, and those introduce injection, exfiltration, and compliance gaps that generic governance frameworks simply don't catch. Continuous monitoring of those vendor AI assets is what's needed, not a one-time assessment done before signing a contract. Third-party risk management teams, InfoSec, Privacy, and Legal all need AI-specific threat detection sitting alongside their existing vendor risk process, because a standard compliance checklist won't surface a real-time injection or exfiltration threat hiding inside a vendor's AI product. This is where a platform built specifically to watch AI-specific threat detection alongside vendor risk assessment for TPRM, InfoSec, Privacy, and Legal teams, since compliance checklists don't surface real-time injection and exfiltration threats in third-party AI assets, fills a gap that traditional third-party risk tooling was never built to cover. PromptArmor, for instance, is an enterprise AI risk intelligence platform focused specifically on continuous monitoring of vendor AI deployments for exactly these risks.

The current vendor landscape: consolidation, acquisitions, and each major platform's coverage

The money flowing into this space gives a decent read on how seriously the market takes the problem. The AI prompt security market reached $1.98 billion in 2025 and is projected to hit $2.61 billion in 2026, a 31.3% compound annual growth rate Research and Markets. Zoom out further and agentic AI security specifically is projected to grow from $1.65 billion in 2026 to $13.52 billion by 2032, a 42.0% CAGR Research and Markets. Those are aggressive growth curves. That's a market that investors and enterprises both believe is about to matter a great deal more than it already does.

2025 was also the year this space consolidated hard, through three acquisitions that reshaped which vendors enterprises can even evaluate as independent options. Protect AI, maker of the Guardian security tooling, was acquired by Palo Alto Networks, with that deal closing in July 2025 and folding into the Prisma AIRS product line. And a third standalone prompt security vendor was acquired by SentinelOne in September 2025 for roughly $159 million in cash and stock.

They're increasingly choosing between large platform suites, folded into bigger security portfolios, and a smaller set of vendors still operating as focused, standalone specialists.

Where does that leave a security team actually making this decision? A broad platform gets you one vendor relationship and one contract, but a specialist vendor sometimes goes deeper on the parts that matter most, in a landscape where 78% of organizations now use AI in at least one business function and the exposed surface is already production-scale, not theoretical BlackFog. Given how fast the threat landscape here keeps moving (340% year-over-year growth in attack volume isn't a number that rewards complacency), that depth question deserves serious consideration rather than defaulting to whichever name is largest OWASP. Lakera, previously the best-known standalone prompt-injection defense vendor (Senthex), was acquired by Check Point (announced September 2025, completed October 2025; reported ~$300M by media; Check Point's FY2025 20-F discloses total consideration of ~$201.8M). Palo Alto Networks Prisma AIRS (v3.0, released March 2026) covers AI application security, AI model security, and AI data protection.

Sources

  1. Enterprise AI Agent Security Solutions: The Complete Buyer's Guide (2026)
  2. Top 5 LLM Security Tools for Enterprise AI Applications in 2026
  3. Prompt injection: types, real-world CVEs, and enterprise defenses
  4. The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism
  5. Indirect Prompt Injection Goes Operational – Lab Space
  6. EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
  7. LLM01:2025 Prompt Injection - OWASP Gen AI Security Project

More in Data Exfiltration