Research
10,000-Trial Taxonomy Maps What Prompt Features Make LLM Agents Exploit Vulnerabilities
Mouzouni presents a systematic taxonomy based on approximately 10,000 trials identifying which system prompt features trigger LLM agents to discover and exploit security vulnerabilities, and which do not. This is the first large-scale empirical mapping of the exploitation surface, moving beyond anecdotal evidence to structured feature attribution. Directly actionable for anyone building agent system prompts — identifies specific prompt patterns that increase or decrease exploit behavior.
Source
↳ Follow the thread