Skills
Context bombs: plant a prompt injection in your AWS secrets and offensive AI agents refuse to continue (57% → 5% admin escalation)
Tracebit found that placing short text designed to trip an attacking LLM's safety guardrails inside canary AWS Secrets Manager values halts autonomous attack agents mid-reconnaissance while alerting the defender that the canary was read. Across 152 runs against Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi K2.6, admin access fell from 57% to 5% (93% → 0% for Opus 4.8), full compromise from 36% to 1%, and any-attack-path completion from 91% to 15%. The technique is free to deploy and doubles as a honeytoken, but Schneier's caveat is the real constraint: it only works against agents that have guardrails, so a locally-run unfiltered model walks straight through it.
Source
↳ Follow the thread