Agents
SADF holds the model constant and finds a 2.6x compromise-rate spread across orchestration frameworks
Julie Brunias presented the Synthetic Agent Deception Framework at DEF CON 34's AI Village — 5,119 evaluation rows, eight architectures, 32 payloads across eight failure modes. Keeping Claude Sonnet fixed and swapping only the wrapper moved Agent Compromise Rate from 11.9% (CrewAI) to 31.1% (SmolAgents), with direct API at 15.5%, LangChain 18.1%, and AutoGen 20.0%. SmolAgents' verbose reasoning traces push intermediate tool output back into the context window, driving 64% Context Boundary Violation and 20% RAG Poisoning — which means model-only safety evaluation systematically misses where the real boundary lives.
Source
↳ Follow the thread