Paper of the Week — Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Coding agents trust rule files blindly — this paper shows attackers can exploit that to trigger package hallucination and slip malicious dependencies into generated code.
Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Yupu Wang, Zhengyuan Jiang, Reachal Wang, Neil Zhenqiang Gong. Published 2026-10-07. arXiv:2610.09272
One sentence summary
Community-shared rule files (AGENTS.md, .cursorrules) are a practical attack surface for prompt injection that reliably causes coding agents to install attacker-chosen packages.
Why this paper
With Claude Fable 5.1, GPT-6 Astra, and Grok 4.6 all deployed in long-horizon agentic coding contexts, rule files have become standard scaffolding — and teams are increasingly pulling them from community repositories. The risk surface is growing fast and this attack vector is almost entirely undefended.
What they did
The authors defined the “package hallucination attack”: an adversary plants malicious instructions inside a shared rule file that nudge the coding agent to import a non-existent or attacker-controlled package name. They evaluated the attack across multiple frontier coding agents and rule-file formats, measuring how reliably the injected dependency name makes it into generated code and gets installed.
Key findings
- Prompt injection via rule files successfully induced package hallucination across all tested coding agents with high reliability, even when the malicious instruction was embedded among otherwise legitimate rules
- Attackers only need write access to a shared rule file on any community platform (GitHub, PromptHub, cursor.directory) — no model access or API keys required
- The attack is durable: once a poisoned rule file is cached or vendored by a team, downstream developers are affected without any further attacker action
- Standard output filtering and safety classifiers did not catch the injected dependency because the output (a plausible package name) looks syntactically valid
- The attack generalizes across AGENTS.md, .cursorrules, and similar formats used by Claude Code, Cursor, and comparable tools
Why it matters for practitioners
If your team pulls rule files from any external source — community repos, internal wikis, shared prompt libraries — you currently have no guaranteed protection against this class of attack. A single poisoned file can silently redirect every pip install or npm add your agent emits toward attacker-controlled packages, and the generated code will look clean on review.
What you can use today
- Treat rule files as untrusted code: add dependency-name allowlists to your CI pipeline so any package name the agent emits is checked against a known-good registry before install
- Pin and audit rule files the same way you pin
requirements.txt— commit them to your repo, diff every update, and don’t auto-pull from upstream without review - If you run Claude Code Auto Mode or similar “primary actor” setups, configure a pre-install hook that surfaces the resolved package source to a human or a separate classifier before execution