Your MCP-Powered Agent Is Already Being Hijacked — 93% of the Time
Your MCP-Powered Agent Is Already Being Hijacked — 93% of the Time
The tool ecosystem your AI agents trust like a package registry has the same security posture as npm circa 2015: anyone can publish, metadata is gospel, and a malicious server can steer your agent wherever it wants. Researchers just automated the entire attack, end-to-end, in a black-box framework that requires zero access to your model weights.
What happened

Researchers from SIAT/CAS built A2M (Attraction-to-Manipulation), a two-stage attack framework targeting agents that use the Model Context Protocol to select and invoke third-party tools. The protocol’s core mechanic — semantic matching between agent goals and tool metadata — is also its core vulnerability: whoever controls the metadata controls what gets invoked. Stage one (“Attraction”) optimizes that metadata to maximize invocation probability; stage two (“Manipulation”) uses execution traces to craft adversarial tool outputs that push the agent toward attacker-chosen outcomes. This is squarely in prompt injection vs. jailbreak territory, but delivered silently through the tool layer rather than through the user-visible conversation. On LiveMCPBench evaluated against GLM-4.6, A2M drove a 93.6% malicious tool invocation rate and a 74.4% mean attack success rate across three attack categories: Information Exfiltration, Environment Integrity Compromise, and Reasoning Derailment. The Cognitive Denial of Service variant inflated token costs to 32.4× the benign baseline — a direct billing attack. Transfer to four other models without any re-optimization still achieved a 63.6% malicious invocation rate and 24.5% attack success rate, suggesting the attack is not model-specific.
Cold read
The 93.6% headline number is optimized and evaluated on the same model (GLM-4.6), which is the most favorable possible condition — the attack was essentially tuned to its target. The cross-model transfer numbers tell the more honest story: 74.4% collapses to 24.5% attack success when you move to unoptimized targets, and the token-cost multiplier drops from 32.4× to 2.7×. That’s still bad, but it’s a different threat level. The benchmark is also a controlled environment — real agentic workflows involve sandboxing, access controls, and monitoring layers that LiveMCPBench doesn’t model. There’s no data here on how the attack fares against MCP deployments that already do tool allowlisting or output validation. Finally, the attack assumes the adversary can publish or compromise a tool server that an agent will actually encounter — a meaningful distribution and discoverability precondition the paper doesn’t cost out.
What it means for you
- Signal maturity: 3/5 — real attack, real numbers, but real-world applicability depends heavily on your deployment’s tool vetting posture
- Who gets hurt: Founders running multi-agent orchestration pipelines that pull from open or third-party MCP registries — especially anyone auto-discovering tools at runtime
- What breaks if this is true: Any MCP-connected agent handling customer data, executing code, or making external API calls becomes an exfiltration or sabotage surface controlled by whoever owns a semantically attractive tool listing
- Why it might not land: Enterprise deployments with static, pre-approved tool registries and network egress controls close the most dangerous attack paths before the framework even gets to run
- Watch for: A CVE or public incident report tied to a named MCP tool registry — that’s when this moves from academic to ops problem
Forecast as of 2026-09-23
By Q2 2027, at least one major MCP registry or orchestration platform (Anthropic, a hyperscaler, or a prominent agentic-AI startup) will ship mandatory cryptographic tool signing or runtime sandboxing as a direct response to this class of attack — or will disclose a breach attributable to semantic supply-chain manipulation of this type.
Source: A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem — Laizhen Li, Xuan Wang, Peicheng Zhao, Juanjuan Zhao, Kejiang Ye, Cheng-zhong Xu, Xitong Gao. https://arxiv.org/abs/2609.26761v1
