IssueTrojanBench: 66.5% of Malicious GitHub Issues Penetrate Every Guardrail in Cursor, Claude Code, and Codex Desktop
Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built a benchmark of malicious issue requests spanning four attack categories and six delivery vectors (PDF attachments, issue comments, and others), then ran it against Cursor, Claude Code, and Codex Desktop backed by GPT-5.3 Codex/GPT-5.4 and Sonnet 4.6. 66.5% of malicious issues cleared both agent-level and LLM-level guardrails, and rejection came almost entirely from the model rather than the agent framework — GPT models were broadly vulnerable while Sonnet 4.6 showed more selective, risk-aware blocking of high-impact actions. For anyone pointing a coding agent at an untrusted issue tracker, this says the framework's own safety layer is close to decorative.
↳ Follow the thread