A recurrent world-model guardrail predicts long-horizon agent risk at 25ms per call instead of judging one action at a time
arXiv 2608.05695·low signal
DreamGuard maintains a recurrent latent state across an agent's trajectory to forecast where a sequence of individually-benign actions is heading, rather than scoring each tool call in isolation. The paper reports the best safety-utility tradeoff among evaluated guardrails at an average 25ms latency per call — cheap enough to sit inline in a production agent loop. This is the defensive counterpart to trajectory-level attacks like Authority-Chain Hijack, which specifically exploit per-action review.