Research
HalluProp Predicts Which Agent Will Fail Before Any Message Is Sent: 84.6% AUROC With 65x Speedup Over Post-Hoc Detection
arXiv 2607.26836 (2026-07-29) attacks cascading failure in LLM multi-agent systems from the pre-hoc side, estimating individual agent failure risk and emergent system-level hallucination risk before inter-agent interaction begins. It models intrinsic risk as fine-grained semantic misalignment between agent role and task query, characterizes propagation via semantic influence plus communication topology, and fuses the two through a differentiable Noisy-OR mechanism. It localizes faulty agents at 84.6% average AUROC with sub-second diagnosis and over 65x speedup versus post-hoc methods — practical as an upstream screen for anyone running fan-out agent orchestration.
↳ Follow the thread