ORCA-bench: Best Frontier Agent Scores 25.3% on Realistic Oncall Root-Cause Analysis, 10.0% on Hard Tasks
ORCA-bench (arXiv 2607.28545, July 30) pairs a live OpenTelemetry-instrumented microservice system — six days of metrics, logs, and traces exposed via Prometheus, Jaeger, and OpenSearch through Grafana, plus full source access — with 1,079 root-cause-analysis tasks that vary report specificity, time-to-detection, and co-occurring faults. Across five frontier agents the best RCA accuracy is 25.3% on Medium (realistic-input) tasks and 10.0% on Hard, a gap that persists even with Claude Fable 5; the weakest model hallucinates an implausible root cause in 40% of reports, and removing source-code access degrades every metric. Ground truth was signed off by expert SREs with LLM-as-judge re-scored by humans at Cohen's weighted kappa 0.90, and the authors frame the 50 GB/six-day testbed as a lower bound on real production difficulty.
Source
↳ Follow the thread