Your AI Agents Are Passing a Virus to Each Other Through Shared Docs

Your AI Agents Are Passing a Virus to Each Other Through Shared Docs

Researchers just demonstrated that a poisoned report can infect one AI assistant, ride into the next artifact it creates, and spread to eight other independent agents — without any human clicking anything malicious. If you’re running a multi-agent workflow where assistants share files, this is your threat model now.

What happened

Figure 7: Second-hop survival S ⁡ ( 2 ) S(2) of candidate templates over the course of the search for the weaker models and DeepSeek-V4-Pro. Bottom axis: cumulative jobs; top axis: cumulative target-m
Figure 7: Second-hop survival S ⁡ ( 2 ) S(2) of candidate templates over the course of the search for the weaker models and DeepSeek-V4-Pro. Bottom axis: cumulative jobs; top axis: cumulative target-m

A team from Cambridge, CISPA, and collaborators built “temporal human-agent universes” — simulated environments where independently operated LLM agents exchange persistent artifacts (think reports, documents, shared files) over time. They studied a specific failure mode they call artifact-mediated propagation: adversarial content enters through a single artifact, gets written into an assistant’s persistent memory, gets reproduced in the next artifact that assistant creates, and then infects the next assistant that reads it — a classic worm pattern, executed entirely through document exchange. This is a subspecies of prompt injection, but the delivery mechanism is the shared artifact layer, not a direct conversation. The results are not subtle: in larger simulated environments, even GPT-5.6 Luna (presumably among the most hardened models available at time of writing) showed 60–80% agent infection rates, with propagation chains reaching eight hops across independent assistants. The attack persisted across extended interaction sequences, meaning it doesn’t die when a session ends.

Cold read

The study uses simulated environments — “temporal human-agent universes” — not production systems, which means the propagation rates (60–80%) reflect controlled conditions where artifact-sharing behavior is essentially guaranteed by design. Real deployments have friction: access controls, sandboxing, human review steps, and memory architectures that vary wildly by vendor. The paper measures whether attacks survive hand-offs, not whether they achieve any particular payload objective — infection rate ≠ damage rate. We don’t know from the abstract what the adversarial content actually does once propagated: exfiltrates data, corrupts outputs, something else? That matters enormously for business risk assessment. Eight-hop propagation is alarming on paper, but in most enterprise multi-agent orchestration setups today, the graph of agents sharing artifacts is not that deep or that connected — yet. Finally, “GPT-5.6 Luna” is not a publicly documented model name as of this writing; it’s unclear whether this refers to a real deployment tier or a research alias, which makes external reproducibility murky.

What it means for you

  • Signal maturity: 3/5 — Real attack vector, simulated blast radius
  • Who gets hurt: Any startup running agentic workflows where multiple AI assistants consume user-generated or third-party documents — legal tech, research tools, enterprise knowledge bases, anything built on RAG pipelines that ingest shared artifacts
  • What breaks if this is true: Your “isolated” agent instances are not isolated if they share a document layer; a single poisoned vendor report or client upload becomes a persistent attack surface that your security perimeter doesn’t even see
  • Why it might not land: Memory architectures differ dramatically — many production deployments don’t have the kind of write-back-to-artifact loop this attack requires; stateless or session-scoped agents are largely immune
  • Watch for: A real-world incident report (not academic) of adversarial content propagating across enterprise agent instances via shared file stores — the moment that happens, every compliance team on earth will require agent memory auditing

Forecast as of 2026-09-29

By Q3 2027, at least one major enterprise SaaS vendor with an agentic AI product will publish a public security advisory or patch specifically addressing artifact-mediated prompt propagation across agent instances — the attack surface is too well-defined and too reproducible for it to stay academic.


Source: Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents — Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav, Vyas Raina, Ivaxi Sheth, Mario Fritz. https://arxiv.org/abs/2609.35576v1

Similar Posts