Fetching from the wire…
Agents2026-08-27 · source-backed
The method synthesizes security skills offline from known attacks and recorded agent failures, injects them into the system prompt at session start, and leaves them active through the tool-use loop. Across six models on RedCode the default all-classes skill dropped malware-generation severity from 3.37 to 0.58 and reached a 43.6% execution attack success rate, comparable to Llama Guard 3's 42% (arXiv 2608.25817). No auxiliary classifier, no execution monitor in the trajectory, which makes it available to API-only deployers who can't touch weights. The paper compares three fixed system-prompt budgets, none of which needs runtime request routing.
Each link below shares sources, entities, or timing with this story.
Shared entity: Llama Guard / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Llama Guard; reported by the same outlet (arxiv.org); overlapping topics (attack, classifier, guard, prompt).
Shared entity: Llama Guard / Same source domain / Shared topic / Earlier coverage
Both cover Llama Guard; reported by the same outlet (arxiv.org); overlapping topics (agent, attack).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (agent, alone, attack, classifier, prompt).
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, alone); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (budget, execution, skill); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (active, attack, budget); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, attack, execution); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (active, agent, alone, budget).