Research
OpsHarness Beats a Bare General Agent on Root-Cause Analysis by 63.4% Without Replacing the Agent
Huang et al. report a finding SREs should note directly: a general-purpose agent such as Codex or Claude Code now often outperforms a purpose-built RCA agent, so the remaining gap lives in the harness rather than the model. OpsHarness pairs layered operational knowledge and an idea-card tool library with a control plane that contrasts successful and failed diagnoses, converts the difference into atomic proposals, and admits updates only through dual-gate verification to block overfitting and regression. It reaches 59.0% top-1 accuracy across two public benchmarks and an industrial deployment, 63.4% above the bare general agent and 4.02x over baseline RCA agents.
↳ Follow the thread