Tool Poisoning Attacks Against LLM Agents
Attackers exploit how AI agents trust tool descriptions without checking if they change mid-session.

Tool poisoning attacks are a form of indirect prompt injection that target the trust an AI agent places in the descriptions of its own tools. The Model Context Protocol made this a live, widespread risk, because its adoption spread faster than anyone built the security review infrastructure to match it.
Tool poisoning attacks and the protocol that made them scalable
MCP was proposed in late 2024. Within less than a year, tens of thousands of MCP servers had gone live, and adoption ran well ahead of any serious security review process. That gap between how fast the protocol spread and how slowly scrutiny caught up is the backdrop for everything that follows.
MCP works in two phases. First, a client connects to a server and pulls down tool descriptions and metadata, the text the agent reads to understand what a tool does and when to use it. Second, once the agent starts working, tool responses flow straight into its context as it runs. That two-step design, description review at connect-time and unchecked response ingestion at runtime, is the structural condition that makes tool poisoning possible in the first place.
The attack itself got a name in 2025, when researchers called it the "Tool Poisoning Attack," or TPA. It's since been formally recognized by OWASP's MCP Top 10 as MCP03, and MITRE ATLAS maps it under AML.T0110 for tool and description poisoning and AML.T0086 for exfiltration via AI agent tool invocation. That's the marker that separates a research curiosity from a named, cataloged threat class that security teams are expected to plan around.
The architectural trust gap that makes tool poisoning work
The vulnerability here is an asymmetry built into how agents handle trust across two different moments in time: descriptions get reviewed once, but runtime content keeps flowing in, unchecked, for as long as the session lasts.
When an agent starts up, it loads tool manifests and reads their descriptions as settled fact, the same ground-truth context it then uses to decide what to do next. This makes the attack so quiet that the agent doesn't need to run a poisoned tool to be compromised. Malicious instructions buried in a tool's metadata get processed the moment the agent plans its next move, before any tool call happens at all.
MCP has no built-in way to catch or block injected instructions hiding in tool descriptions, parameter schemas, or response content. Attackers lean on that gap with plain concealment tricks, zero-width space characters, or adversarial text tacked onto the end of an otherwise normal-looking description, so that anyone doing a manual review sees nothing wrong.
A newer variant called WebMCP pushes the problem further. Research published as arXiv:2606.06387 (Lee et al.) shows that in WebMCP, the set of available tools isn't fixed even within a single session: third-party scripts can inject or swap tools mid-session, a technique the researchers call Mid-Session Tool Injection. That breaks the one assumption defenders had left, that whatever tool set got reviewed at connect-time would stay put. If the tools themselves can change while the agent is mid-task, a one-time check at the door stops meaning much.
The three main attack variants and the trust gap they operationalize
Three distinct attack patterns exploit that same trust gap, each at a different point in a tool's life. That's why no single, point-in-time control shuts down all three at once.
Description poisoning is the simplest. Malicious directives get written straight into a tool's metadata at registration, and they're live for the entire session from the moment the manifest loads. No second step, no trigger needed.
Rug-pull attacks work on a longer clock. A tool passes review looking completely clean, earns real trust over days or weeks of normal use, and only later has malicious logic slipped in. Researchers demonstrated this against a WhatsApp MCP setup in April 2025: a secondary server first advertised a harmless "fact of the day" tool, then swapped its description for a malicious one on the second load, instructing the LLM to exfiltrate the user's entire WhatsApp message history through what looked like an unremarkable tool call, with nothing in the output to tip anyone off. A real-world case followed the same pattern in the open-source supply chain. The postmark-mcp npm package shipped clean through versions 1.0.0 to 1.0.15. Version 1.0.16, released September 17, 2025, added one line of code that BCC'd every outgoing email to an address the attacker controlled. Snyk confirmed how the mechanism worked, and Koi Security estimated roughly 300 organizations were affected before the package got pulled, with several thousand emails a day flowing to the attacker's domain at peak. Scanning a tool once at registration tells you nothing about what it looks like three weeks later.
Tool shadowing closes the loop. Here, a poisoned description doesn't attack its own tool directly. It changes how the agent uses a different, trusted tool, so the agent's own trust in a clean server becomes the path data takes on its way out, crossing server boundaries along the way.
Each variant answers the naive fix for the one before it. Scan at registration? Rug-pulls show that's not enough, since the malicious code is added after review has already passed. Monitor tools individually? Shadowing shows a clean tool can still become the exfiltration path once a poisoned neighbor redirects it.
The attack surface's expansion into supply chains, lifecycle hooks, and persistent agent state
The same root cause, agents treating externally supplied content as trustworthy instructions, has spread well past MCP tool descriptions. It now shows up in skill marketplaces, lifecycle hooks, and an agent's own persistent memory.
On the supply-chain side, researchers describe a technique called Document-Driven Implicit Payload Execution, or DDIPE, where malicious logic hides inside the code examples and configuration templates found in ordinary skill documentation. An agent following that documentation reproduces the payload as part of doing its normal job, with no explicit prompt telling it to. The scale here is substantial: the public SkillsMP marketplace lists over 631,813 skills, and none of them face mandatory security review. CVE-2025-59536 showed what that looks like in practice, attackers planted hook-laden configuration files in repositories that bypassed the startup trust dialog entirely, achieving remote code execution and pulling API keys. A separate flaw in Smithery's build pipeline, a path traversal bug disclosed in October 2025 and patched back in June 2025, exposed a single overprivileged Fly.io API token that could have given theoretical control over a large share of hosted MCP apps, though no one found evidence it was exploited in the wild.
Lifecycle hooks open a different door. These are configuration entries that bind shell commands to events like session start, a tool call, or a file edit, and researchers (Li et al., arXiv:2609.03884) call this out as its own attack surface, one they term HOOKPRY. Hooks run as subprocesses on the host machine, entirely outside the LLM's own reasoning loop, so the model never sees or approves the command and no prompt-level defense can inspect it. A plugin that's perfectly safe today can be trojanized by a later update that quietly binds an attacker's command to a benign-looking event, a rug-pull one layer below where most people are looking. HOOKPRY was tested against harnesses including Claude Code, Codex CLI, and OpenCode, and across every combination tried, it compromised all seven harnesses evaluated. Microsoft Defender caught none of it, zero recall, and even combining three separate static defenses together still missed close to half of the malicious artifacts.
Then there's the attack on an agent's own memory. Researchers (Li et al., arXiv:2605.28201, 2026) formalize what they call the Sleeper Attack: adversarial content gets planted in an agent's session context, memory, or a reusable skill, sits dormant, and only activates later when a completely ordinary user query triggers it. Because the planting and the payoff happen in two separate interactions, tracing the attack back to its source is much harder than with a single-interaction exploit. The paper lays out three distinct strategies for doing this, which it names Latent Instruction Planting, Proactive Information Elicitation, and Persistent Information Corruption. The pattern across all three expansions, supply chain, lifecycle hooks, persistent memory, is the same: the attack surface is migrating outward faster than any one defense can follow it.
Why capable models are more exploitable, not less
A reasonable instinct says a smarter model should be harder to fool. The evidence points the other way. The instruction-following skill that makes a frontier model genuinely useful is the same skill that makes it reliably exploitable through tool poisoning, since a stronger model follows poisoned instructions more faithfully, not less.
The MCPTox benchmark (Wang et al., 2025; arXiv:2508.14925) tested a range of prominent LLM agents and found o1-mini, a model picked specifically for its strong instruction-following, recording a 72.8% attack success rate. That number lands on the model's core strength, not a weakness around its edges. A model built to follow instructions closely will follow poisoned ones embedded in trusted tool context just as closely.
This isn't a problem unique to MCP, either. The Sleeper Attack benchmark shows agents that score low on attack success under ordinary, single-interaction tests remain just as vulnerable once the attack plays out across multiple interactions over time. Capability in the moment doesn't translate into resistance across a longer time horizon. If anything, the more an agent is designed to trust and act on what it's told, the more useful that same design becomes to whoever controls what it's told.
Where the current defensive ecosystem falls short
A real body of defensive work has emerged, and each proposal makes a genuine dent, but each one covers a slice of the problem, not the whole connect-time-to-runtime gap, and each comes with its own cost in speed or usability.
VIGIL (arXiv:2601.05755, 2025) intercepts tool stream injections with a verify-before-commit step before the agent acts, which handles runtime injection but says nothing about metadata that was poisoned before the session even started. TRUSTDESC (arXiv:2604.07536, 2026) takes the opposite angle, generating independently verified tool descriptions to replace ones an attacker controls, directly addressing description poisoning, but only if an organization has the trusted verification infrastructure to run it. (2026), quarantines untrusted descriptions during an isolated first planning pass so poisoned metadata can't steer the agent's actions, and it holds up well on AgentDojo and Agent Security Bench without sacrificing task usefulness, though it adds latency to every run.
MCP-Guard (arXiv:2508.10991, 2025) sits as a proxy doing real-time filtering with lightweight, pattern-based static scanning, which is fast, but scanning methods at this stage still carry both high false-negative and high false-positive rates at once, meaning real attacks slip through while legitimate workflows get blocked. Zero-trust registries, where an administrator controls registration and tools carry a dynamic trust score, cut down on tool squatting and poisoning but bring their own latency and ongoing maintenance burden, especially in environments where the tool set changes often. HARD (Ruan et al., arXiv:2608.12977) takes a different approach entirely, a self-evolving runtime framework that improves its own defenses based on failures it's actually observed, a shift away from static, handwritten rules toward something that adapts on its own, though it needs real operational infrastructure running behind it to work.
Publicly available scanning tools still struggle to reliably tell malicious tools from safe ones, landing in that same bind of high false negatives alongside high false positives. The risk extends beyond the application layer. CVE-2025-6514, disclosed in July 2025, hit mcp-remote with a CVSS score of 9.6 for OS command injection triggered when the proxy connected to an untrusted server, in a package that had been downloaded more than 437,000 times. Infrastructure-level tooling carries this exposure just as much as any individual agent does.
The strongest objection to the entire defensive conversation is this: the root cause is architectural. Any system where an LLM reads externally supplied natural-language metadata and acts on it autonomously reproduces this same vulnerability, no matter what protocol sits underneath. Fixes built specifically for MCP matter and are worth building, but they can't be the whole answer, because the problem was never really about MCP's particular wiring. It's about what happens any time a model is handed instructions it can't tell apart from its own operating assumptions.
Static governance reviews versus a dynamic, runtime attack surface
Traditional third-party risk management, the compliance checklist a vendor fills out once before onboarding, assesses a tool at a single point in time. That's precisely the snapshot this whole attack class is built to survive.
Rug-pull attacks and the trojanized lifecycle hooks described earlier both make the same point from different angles: a tool can look completely clean during review and turn malicious days or weeks afterward. Pre-deployment scanning simply can't see that far ahead, because the change hasn't happened yet when the scan runs.
Frameworks like the OWASP MCP Top 10 and MITRE ATLAS give the field a shared vocabulary for naming these threats, and that matters. But naming a threat after the fact isn't the same as catching it in the moment. Those frameworks classify what already happened; they don't catch a poisoned tool description the instant it appears inside a live agent session.
A basic category error sits underneath a lot of generic governance paperwork: it treats tool descriptions as documentation to file away, not as executable instructions a model will act on, and that mistreatment is what leaves the gap the attack exploits. That's the exact gap the attack is built to exploit. And the WebMCP research on Mid-Session Tool Injection adds one more wrinkle, the tool surface itself can change mid-session, well after whatever point-in-time review already happened. A review that only ever looks at the start of a session has nothing to say about what the session becomes an hour in.
For the teams actually responsible for vendor AI risk, third-party risk management, information security, privacy, and legal, the shape of the problem doesn't match the shape of the review process built to catch it, and the implication is that vendor approval can't remain a one-time gate. A checklist completed once at onboarding can't watch a tool that keeps changing after onboarding ends.
Sources
- Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
- A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- WebMCP Tool Surface Poisoning: Runtime Manipulation Attacks on LLM Agents
- MalTool: Malicious Tool Attacks on LLM Agents
- TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation
- VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit


