Network Troubleshooting Agents Are Near-Saturated on Real Faults but Invent Root Causes When Nothing Is Broken
FaulT-Bench runs 200 troubleshooting scenarios across eight network topologies, five reimplemented from public practitioner labs, deployed in Kathara and spanning genuine faults, false fault reports, wrong device attribution and wrong root-cause claims. Evaluating SADE, ReAct and Claude Code, all three are near-saturated on accurate tickets and robust to misdirection but degrade sharply when the network is healthy and the ticket is wrong, probing until a benign condition can be promoted to a root cause instead of concluding nothing is wrong. Rewriting 72 false-premise tickets into five reporter personas shows wording matters more than content: a confidently wrong report is handled about as well as an accurate one, while a vague underspecified one collapses performance.
↳ Follow the thread