Agents
HARD: let the agent evolve its own runtime defenses instead of hand-writing guardrails
arXiv 2608.12977 (2026-08-13) makes the case that handcrafted agent defenses are structurally behind — each new attack class needs a human to write a new rule — and proposes HARD, a framework where the defense itself goes through an autonomous evolution loop. Reported results improve security over existing handcrafted defenses while preserving benign task utility. Read alongside the same week's skill-misevolution paper it is a pointed pairing: self-evolution is being proposed as the fix for defenses in one paper and identified as the vulnerability in agents in the other.
Source
↳ Follow the thread