Research
DiagChain: Best LLM Agent Configuration Reconstructs Only 39.6% of 849 Attack-Chain Steps, With Failures Splitting by Model Size
DiagChain evaluates security agents stage-by-stage rather than on final output, using MAIN-69 — 69 scenarios across multiple operating systems, evidence noise levels, and chain lengths — plus an Evidence-Centric RAG method that couples retrieval to an evolving structured chain representation. Across six LLMs, the strongest configuration succeeds on just 39.6% of 849 reference steps. The diagnostic split is the useful part: smaller models fail at the basic task of incorporating retrieved evidence into their output, while larger models get further and then bottleneck on correctly ordering that evidence.
↳ Follow the thread