Research
HARD Turns LLM Agent Runtime Defense Into a Self-Evolving Loop Instead of Hand-Written Rules
The authors give a harness-level formulation of runtime defense for LLM agents, characterizing how harness mechanisms enable defense construction and unifying existing interventions under one view. Building on it, HARD (Harness-based Autonomous Runtime Defense Evolution) automatically picks intervention strategies and iteratively improves defense artifacts from observed failure traces, reporting better security than handcrafted defenses while preserving benign task utility. The framing matters more than the numbers here: it treats the agent harness — not the model — as the defensible surface, and makes failure traces the training signal for the guardrails themselves.
↳ Follow the thread