The AI App Store Has a Malware Problem Nobody Can Detect
The AI App Store Has a Malware Problem Nobody Can Detect
You’re installing third-party skills into your AI agents the same way you once installed browser extensions — and we all remember how that ended. A new paper just showed that the best available scanning defense gets beaten 97% of the time by an attacker who knows how it works. Your agent’s “skill marketplace” is an open wound.
What happened
Researchers Kaisar and Dhar built Pretext, a white-box adversarial tool that crafts malicious agentic workflow skills — the plugin-style instruction packages used by agents like OpenClaw and Claude Code — so they sail past automated security scanners. The target defense architecture pairs deterministic static analysis with an LLM-as-judge semantic reviewer (as in NVIDIA’s SkillSpector). Pretext defeats it with two moves: shift the payload from code into natural language (killing the static-analysis signal), then frame it as legitimate purpose and split instructions across multiple files so the LLM stage never sees enough in one pass to trigger its block threshold. Against a frozen (non-adapting) detector, Pretext hit 97% evasion; against a co-adaptive detector that fights back, it still reached 77%. All results are across three open-source models.
Cold read
The 97% number is a white-box result — the attacker has full knowledge of the detector’s architecture and weights. In the real world, most attackers do not get that. The 77% co-adaptive figure is more operationally honest, but “co-adaptive” is also a lab construct; real defenders iterate on a calendar of weeks or months, not adversarial training loops. The paper tests three open-source models, not the closed frontier models (GPT-4o, Claude 3.x) that actually power the highest-value enterprise agents — transferability is unproven. And the benchmark is evasion rate, not real-world damage rate: a skill that evades detection but fails to execute its payload in a live environment is not actually dangerous. None of this is debunked by the paper; it’s just outside what the abstract claims.
What it means for you
- Signal maturity: 3/5 — attack is real and well-constructed; defense applicability to production stacks is still unclear
- Who gets hurt: Founders running multi-agent platforms with third-party skill/plugin marketplaces — think AI dev-tool companies, enterprise automation builders, anyone with an “extend your agent” feature
- What breaks if this is true: The entire “scan before install” trust model for agent skills collapses, meaning you cannot outsource security judgment to a scanner and go back to humans-in-the-loop or allowlist-only skill registries
- Why it might not land: White-box evasion against closed, continuously updated detectors is materially harder; if the big providers (Anthropic, OpenAI) treat their judge-model as a secret and retrain it frequently, the attack surface shrinks significantly
- Watch for: A real-world incident — compromised skill in a public marketplace executing unauthorized actions in a production agentic workflow — that forces a platform to pull its third-party skill registry; that’s the moment this goes from research to policy
Forecast as of 2026-10-01
By Q3 2027, at least one major agent platform (OpenAI GPTs, Anthropic’s tooling, or a top-5 enterprise AI automation vendor) will publicly tighten or suspend open third-party skill submission in direct response to demonstrated marketplace evasion — either citing this research or a live incident it foreshadowed.
Source: Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents — Tobias Kaisar, Aritra Dhar. https://arxiv.org/abs/2609.39607v1
