Fetching from the wire…
Agents2026-08-27 · source-backed
The target is attacks where a harmful objective is split across individually plausible requests and tool calls, so the harm is only visible in the accumulated trajectory. Existing defenses either pay for auxiliary online reasoning or judge actions after generation, which ties them to a specific runtime action representation (arXiv 2608.25711). ReDiR injects a compact latent safety representation into the frozen base model at generation time, learned via same-model cross-view supervision, and holds attack success below 8% across two agent-safety benchmarks, three model families and eight held-out tool domains.
Each link below shares sources, entities, or timing with this story.
Shared entity: Existing / Same source domain / Shared topic / Earlier coverage
Both cover Existing; reported by the same outlet (arxiv.org); overlapping topics (attack, model).
Both cover Existing; reported by the same outlet (arxiv.org); overlapping topics (latent, model).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (attack, model, safety, tool); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (action, attack, call, tool, trajectory).
Reported by the same outlet (arxiv.org); overlapping topics (action, attack, call, model, tool).
Reported by the same outlet (arxiv.org); overlapping topics (action, benchmark, call, model, tool).
Shared entity: Existing / Same source domain / Earlier coverage / Tension
Both cover Existing; reported by the same outlet (arxiv.org); earlier Existing coverage from 2026-07-21.
Both cover Existing; reported by the same outlet (arxiv.org); earlier Existing coverage from 2026-07-17.