Research
Recast Forecasts Multi-Turn Jailbreaks 2.41 Turns Ahead, Catching 88.3% of Future Safety Failures
Recast (arXiv 2607.26820, 2026-07-29) moves LLM safeguarding from turn-level violation detection to trajectory-level risk prediction, on the premise that malicious intent gets decomposed across individually harmless turns and reassembled later. It retrieves risk-relevant evidence from both short-term dialogue progression and long-term history via a dual-scale trajectory view, then a causal temporal encoder predicts the distribution of future risk-emergence turns. Across 7 risk categories it predicts 88.3% of future safety failures with an average 2.41-turn lead time at a 12.3% false alarm rate — enough headroom to intervene rather than react.
↳ Follow the thread