Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.09885 treats the harness as the thing that evolves with emerging risk rather than a static wrapper around a model you keep re-aligning. Four artifacts with non-overlapping responsibilities: System Prompt, Rule Bank, Safety Memory, Tool Policy. Failures get attributed to one artifact and fixed locally. It beat a static SafeHarness baseline 3.1x on Agent-SafetyBench while improving utility, generalized to unseen risks on AgentHarm, and transferred across models without retraining. This is the pattern I'd build against if I were designing agent safety from scratch today.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage / Tension
Both cover Agent, SafetyBench; reported by the same outlet (arxiv.org); earlier Agent coverage from 2026-08-10.
Shared entity: Agent / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Agent; reported by the same outlet (arxiv.org); overlapping topics (agent, beat).
Both cover Agent; reported by the same outlet (arxiv.org); overlapping topics (agent, model).
Shared entity: SHE / Same source / Shared topic
Both cover SHE; cite the same source (arXiv 2608.09885); overlapping topics (agent, safety).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, baseline, beat, model); pushes against this story (against).
Shared entity: Agent / Shared topic / Earlier coverage / Tension
Both cover Agent; overlapping topics (against, agent, around); earlier Agent coverage from 2026-08-03.
Shared entity: Agent / Same source domain / Shared topic / Earlier coverage
Both cover Agent; reported by the same outlet (arxiv.org); overlapping topics (against, agent).
Both cover Agent; reported by the same outlet (arxiv.org); overlapping topics (against, agent).