LLM Security Review

Direct vs Indirect Prompt Injection Explained

How LLMs confuse instructions from data in two dangerous ways.

Staff Writer · · 13 min read
Cover illustration for “Direct vs Indirect Prompt Injection Explained”
Prompt Injection · September 29, 2026 · 13 min read · 2,950 words

Direct vs Indirect Prompt Injection Explained. Direct and indirect prompt injection share the same root vulnerability (the LLM's inability to separate instructions from data), but differ fundamentally in attack surface, attacker position, and real-world risk, and understanding that distinction is the foundation for any coherent defense strategy.

Why LLMs cannot separate instructions from data

A large language model processes developer instructions and untrusted content as a single stream of tokens, with no privileged channel that separates "instruction" from "data" Forcepoint X-Labs. Whatever text ends up in the context window can act as a command, regardless of where it came from or who put it there.

Researchers call this the semantic gap. It's a useful comparison to SQL injection, where an attacker sneaks executable code into what's supposed to be a plain data field. But SQL injection has a fix that actually works: parameterized queries separate code from data at the database layer, and the problem mostly goes away. Prompt injection has no equivalent. Natural language doesn't come with a schema, and it can't be validated like a form field, because the whole point of natural language is that it's flexible, ambiguous, and context-dependent.

Prompt injection and jailbreaking get lumped together constantly, but they differ in a key way. Jailbreaking targets a model's built-in safety training, the stuff that's supposed to stop it from generating harmful content. Prompt injection targets something more basic: the model's inability to tell instructions apart from data. A jailbreak might get a model to describe something it shouldn't. A prompt injection can get a model to run a command it was never authorized to run.

None of this is a bug sitting in a queue waiting for a patch. It's a structural property of how transformer-based language models work, and that's precisely why OWASP lists prompt injection as the number one risk in its Top 10 for LLM Applications Forcepoint OWASP GenAI Security Project Q1 2026 Exploit Round-up Report.

The mechanics of direct prompt injection

An attacker has access to the input field, whether that's a chatbot window, a support widget, or an internal copilot, and types the malicious instruction straight in. No intermediary, no waiting.

Every direct injection shares two structural components. First, a trigger phrase designed to make the model drop its system instructions: something like "ignore all previous instructions" or "you are a new assistant now." Second, the actual payload, whatever the attacker actually wants: data exfiltration, an instruction override, a persona switch. A third layer occurs often enough to count as a pattern rather than an exception: obfuscation, using Base64 encoding, Unicode tricks, or mixing in other languages, all aimed at slipping past content classifiers that are watching for known bad phrases Forcepoint X-Labs.

Researchers Perez and Ribeiro identified two subcategories back in 2022 that still hold up. Goal hijacking forces the model to execute the attacker's task instead of whatever the developer intended OWASP GenAI Security Project Q1 2026 Exploit Round-up Report Forcepoint X-Labs. Prompt leaking goes after something different: extracting the hidden system prompt or confidential context the developer never meant to expose. Jailbreaking, in this framework, is really a specialized subset of direct injection, aimed specifically at bypassing safety filters rather than hijacking the task itself.

The clearest early proof that this doesn't require any technical sophistication came in February 2023, when a Stanford student got Microsoft's Bing Chat to reveal its internal codename, "Sydney," along with its full system prompt. No exploit kit, no elevated access, just carefully worded natural language. That's the thing about direct injection: the instruction is sitting right there in the chat log. Anyone reviewing the session can see it, which makes direct injection comparatively easier to catch than what comes next.

Why indirect prompt injection is structurally harder to stop

Indirect injection flips the attacker's position. The attacker never touches the AI system at all Forcepoint X-Labs. Instead, the malicious instruction gets buried inside content the model reads on behalf of a user, emails, PDFs, web pages, documents, API responses, retrieved files Forcepoint X-Labs. The user who triggers the interaction sees nothing unusual. They just asked their assistant to summarize an email. The model is the one that stumbles onto the hidden instruction and treats it as a legitimate command.

Greshake and colleagues formalized this in 2023, extending the earlier Perez and Ribeiro threat model to cover adversarial content embedded in third-party documents Forcepoint X-Labs. Liu and colleagues followed up in 2024 with a benchmark of defenses Forcepoint X-Labs.

One variant deserves its own mention: stored injection. A malicious payload gets written into a database, a memory store, or some persistent document, and it sits there until a future interaction pulls it back up Forcepoint X-Labs. The attacker's instructions outlive the session that planted them. That's a genuinely different threat model than a one-off chat message; it means an attack from three weeks ago can still be live today.

The obfuscation techniques here are more elaborate than the ones used in direct injection, mostly because the target isn't a content filter watching a chat box; it's a human never even looking at the raw markup. CSS-suppressed text, zero-pixel fonts, HTML comments, meta-tag namespaces: all invisible to someone scanning a rendered web page, all fully visible to an LLM parsing the raw DOM. Character encoding tricks and hidden Unicode tokens are meant to slide past keyword classifiers, producing a payload that a security team could stare directly at and never notice.

That's the core of why indirect injection is so much harder to catch. The malicious instruction never appears in the user's input, so any monitoring built around screening what users type simply won't see it. It arrives through a content pipeline, an email inbox, a crawled webpage, a retrieved document, that most security teams aren't watching with anything like the same rigor Forcepoint X-Labs. And the economics tilt hard toward the attacker: a defender has to screen every document, every page, every email, every API response the agent might touch, while the attacker only needs one payload that works once OWASP GenAI Security Project Q1 2026 Exploit Round-up Report. That asymmetry doesn't favor patience. It favors volume.

Direct and indirect injection across the dimensions that matter for defense

Start with attacker position. Direct injection requires the attacker to actually have access to the input interface, a live user session, an exposed API endpoint. Indirect injection requires nothing of the sort. The attacker just needs to place content somewhere the model will eventually read it, a public web page, a shared document, a poisoned email sitting in someone else's inbox OWASP GenAI Security Project Q1 2026 Exploit Round-up Report Forcepoint X-Labs. One requires a foothold. The other requires patience and a hosting account.

That difference in position drives a difference in scope. Direct injection is bounded by wherever the model happens to accept user input, which is a known, enumerable surface. Indirect injection is unbounded: every document, every web page, every RAG-indexed file, every tool output the agent touches becomes a potential delivery vehicle Forcepoint X-Labs. A security team can list the input fields on their chatbot. They cannot list every web page an autonomous browsing agent might visit next week.

There's also the question of who's actually driving the interaction. In direct injection, the attacker is the one typing, the one present in the conversation. In indirect injection, the legitimate user is the one who unknowingly triggers the payload; the attacker doesn't need to be anywhere near the system once the content's been planted.

Visibility follows from that. A direct injection attempt is in the session log, plain text, ready for a classifier or a human reviewer to catch. Catching an indirect payload requires monitoring the content pipeline itself, since it lives inside external content and may never surface for a human at all Forcepoint X-Labs.

And the consequences scale differently too. Direct injection tends to be session-scoped: embarrassing outputs, a leaked system prompt, a bypassed safety filter, contained mostly to that one exchange. Indirect injection can persist across sessions through the stored variant, trigger with zero clicks from the victim, and cascade into tool calls, data exfiltration, and lateral movement across whatever permissions the agent happens to hold Forcepoint X-Labs. Same root cause. Very different blast radius.

The "lethal trifecta" that turns any agent into a viable target

Independent researcher Simon Willison put a name on the pattern in 2025 that ties most of this together: the lethal trifecta. Three properties have to show up at once for an agent to be exploitable through indirect injection. Access to private data. Exposure to untrusted content Forcepoint X-Labs. And the ability to communicate externally or take some consequential action.

Missing any one leg of that triangle breaks the attack path OWASP GenAI Security Project Q1 2026 Exploit Round-up Report. An agent that reads untrusted web pages but has no private data to leak and no way to send anything out isn't much of a target. An agent that holds sensitive data but never reads anything outside a locked-down internal system has no delivery mechanism for a payload to arrive through. It's only when all three legs are present, together, in the same agent, that the exploit chain actually closes.

That framework explains why tool-call hijacking became the dominant failure mode heading into 2026 Forcepoint OWASP GenAI Security Project Q1 2026 Exploit Round-up Report. Picture an agent with a send_email tool, reading a document that contains an injected instruction to redirect the email to an attacker-controlled address Forcepoint X-Labs. The model doesn't flag anything as unusual. It just picks the wrong recipient and sends the message, because from its perspective, the instruction embedded in the document reads exactly like a legitimate one.

The underlying vulnerability didn't change between the early jailbreak demonstrations and now; what changed is what agents are allowed to do. Retrieval, tool use, MCP integrations, the whole agentic shift, took indirect injection from a theoretical concern in a research paper to the dominant attack pattern in production systems. More capability gives the trifecta more surface to complete itself on.

Anthropic's Claude Opus 4.5 System Card, published in November 2025, put numbers on how this compounds Forcepoint. Using Gray Swan's Shade tool in agentic coding environments, a single strong indirect prompt-injection attempt succeeded 4.7% of the time Forcepoint. Ten attempts from the attacker push that to 33.6% Forcepoint. At a hundred attempts, it reaches 63.0% Forcepoint.

Diagram: How Repeated Injection Attempts Compound Success Rates. Visualizes: Show how the probability of a successful indirect prompt injection rises sharply as an attacker makes repeated attempts, using the three data points from Anthropic's…

Five production incidents from 2025–2026: real-world exploitation

EchoLeak is probably the clearest case study on record for what a fully chained indirect injection attack looks like in production. The chain stacked multiple bypasses on top of each other, evading Microsoft's XPIA classifier, getting around link redaction using reference-style Markdown, exploiting auto-fetched images, and abusing a Teams proxy the content security policy happened to allow Forcepoint X-Labs. It's described as the first known case of prompt injection weaponized into concrete data exfiltration against a live production system. Microsoft shipped a server-side patch in May 2026, though the underlying category of risk, RAG-based assistants reading untrusted content, doesn't go away just because one path got closed.

Palo Alto's Unit 42 documented a different flavor of the same problem: the first reported real-world detection of an indirect injection built specifically to bypass an AI-based ad review system. What stood out wasn't just that it worked, it's that it stacked multiple indirect injection methods at once, a level of sophistication beyond earlier single-technique detections. Unit 42's broader research mapped twenty-two distinct payload-delivery techniques already active in the wild Palo Alto Networks Unit 42.

Cursor fixed it in version 1.3.9; every earlier release stayed exposed. The victim didn't need to do anything unusual, just an ordinary prompt that happened to pull in attacker-controlled content from an MCP server response or a poisoned search result OWASP GenAI Security Project Q1 2026 Exploit Round-up Report Forcepoint X-Labs. Given that Cursor is used by more than half of the Fortune 500 by its own count, the exposure window mattered. Cursor 3.0, released April 2, 2026, closed it; every version before that stayed vulnerable.

The fifth case is less a single incident than a pattern. Forcepoint X-Labs confirmed ten verified indirect injection payloads deployed on real, live websites in 2026, not lab demonstrations, engineered to activate only when an AI agent reads the page. The targets ranged across financial fraud, data destruction, API key theft, and denial-of-service against agents themselves; one payload embedded a fully specified $5,000 PayPal transaction request, clearly built for a browser agent or financial assistant holding stored payment credentials OWASP GenAI Security Project Q1 2026 Exploit Round-up Report Palo Alto Networks Unit 42 and Forcepoint. CVE-2025-32711 carries a CVSS score of 9.3 and was disclosed by Aim Security researchers. In this zero-click indirect prompt injection, an attacker sends a crafted email that the user never opens, Copilot reads it during background processing, and a later unrelated query triggers exfiltration of internal files to an attacker-controlled server. CurXecute (Cursor IDE, 2025). CVE-2025-54135 has a severity of 8.6 (catonetworks.com, Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives). DuneSlide (Cursor IDE, disclosed July 2026). CVE-2026-50548 and CVE-2026-50549, both carrying a CVSS score of 9.8, were discovered by Cato AI Labs (catonetworks.com, Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives). Google's crawl data (roughly 2–3 billion pages per month) confirmed a 32% relative increase in malicious indirect prompt injection content between November 2025 and February 2026 (Forcepoint, Google Security blog). A cross-cutting observation from the OWASP GenAI Q1 2026 Exploit Round-up (published April 14, 2026) found that of eight major AI-related incidents from January through April 11, only one received a CVE identifier, with the rest stemming from misconfiguration, excessive agency, supply-chain failure, or prompt injection, underscoring that standard vulnerability tracking infrastructure is not keeping pace with AI-specific exploits (OWASP GenAI Security Project Q1 2026 Exploit Round-up Report).

Why no single defense closes the gap

OpenAI has called prompt injection a frontier security problem that's still being worked through, not solved. That framing got official backing in May 2026, when the Five Eyes intelligence alliance, CISA and the NSA alongside counterparts in the UK, Canada, Australia, and New Zealand, issued joint guidance naming prompt injection a core manipulation vector and stating that no single safeguard is sufficient.

Input filters and classifiers, tools like Azure's Prompt Shields or Llama Guard 3, catch known attack patterns reliably enough Splunk Forcepoint. But adaptive attacks using encoding tricks, multilingual phrasing, or payloads split across multiple messages slip through regularly; one 2025 challenge found that 10% of more than 300,000 attempts succeeded against basic safety filters Splunk Forcepoint. Spotlighting, a technique Microsoft Research introduced in 2024, marks untrusted text with special delimiter tokens so the model can tell it apart from trusted instructions Forcepoint X-Labs. It helps, but it does nothing for retrieved content the model was never told to treat as untrusted in the first place Forcepoint X-Labs. Prompt sandwiching, repeating the trusted instruction after the untrusted content so the model doesn't lose track of it, catches some goal-hijacking attempts but does nothing against stored or zero-click variants.

Simon Willison's dual-LLM pattern, proposed in 2023, is architecturally sound in principle: a privileged model handles trusted instructions, a quarantined model handles external content, though in practice it adds latency, complexity, and its own integration surface that can itself be attacked Forcepoint X-Labs.

The same asymmetry mentioned earlier drives all of these: defenders have to screen every channel an agent touches, while an attacker only needs one payload that gets through any single one of them OWASP GenAI Security Project Q1 2026 Exploit Round-up Report. Sysdig's analysis put it bluntly: adaptive attacks bypass essentially every published defense when that defense is tested in isolation. The research community is fairly candid that no single technique closes the gap on its own.

What a coherent defense requires across three layers

So where does that leave an organization actually trying to ship an agent safely? Start with architecture, before any attack ever arrives. Run every agent through the lethal trifecta as a design checklist: does this agent need access to private data, exposure to untrusted content, and external communication? If any one leg isn't functionally required for the job the agent's supposed to do, strip it out Forcepoint X-Labs. An agent that summarizes internal documents doesn't need outbound email access. An agent that drafts emails doesn't need to browse arbitrary web pages. Cutting one leg off the trifecta before deployment does more for security than any classifier bolted on afterward, because it shrinks what's exploitable rather than trying to catch the exploit after the fact.

Layered on top of that architectural discipline sit the detection and containment measures already covered: input classifiers for what they're worth, spotlighting to help models flag untrusted spans, sandwiching to reinforce the original instruction, and dual-LLM separation where the latency cost is tolerable. None of these substitute for the trifecta check. They're a second line, catching what slips past a design that's already been narrowed down as far as the agent's actual job allows.

The last layer is less technical and more organizational: monitoring the content pipeline itself, including the chat log. Since indirect injection routes through emails, documents, and API responses a security team may never think to watch, closing that blind spot means treating retrieved content with the same scrutiny given to user input, logging what an agent reads, not only what it's told Forcepoint X-Labs. None of the three layers, architecture, technical defenses, pipeline monitoring, does the whole job alone. Together, they're the closest thing available right now to a defense that actually holds.

Sources

  1. The Prompt Injection Risk No Single Control Can Stop
  2. The Comprehensive Guide to Prompt Injection Attacks in 2026 | Sysdig
  3. What Is Prompt Injection? Understanding Direct Vs. Indirect Attacks on AI Language Models | Splunk
  4. Indirect Prompt Injection Goes Operational
  5. Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives
  6. LLM01:2025 Prompt Injection
  7. simonwillison.net
  8. catonetworks.com
Filed underPrompt Injection

More in Prompt Injection