Fetching from the wire…
Public story · 2026-07-31 · high
ORCA-bench found the weakest of five frontier agents invented oncall root causes 40% of the time, using a live incident system with full source access.
Why now: The benchmark's numbers landed in the July 31 briefing, as teams weigh putting agents on oncall rotations.
The best of five frontier agents solved just 25.3% of realistic root-cause tasks on a new oncall benchmark, per ORCA-bench's authors. That's the number SRE and platform teams should plan around before handing an agent the pager. On the Hard tier, the top score drops to 10%. The weakest model invents an implausible root cause in 4 of every 10 reports.
ORCA-bench runs on a live, OpenTelemetry-instrumented microservice system. It provides six days of real metrics, logs and traces from Prometheus, Jaeger and OpenSearch, plus full access to the source code. Agents work through 1,079 tasks that vary how specific the incident report is and whether more than one fault is firing at once. Ground truth came from expert SREs, with agreement scored at a Cohen's weighted kappa of 0.90, so the scoring isn't the weak point.
Claude Fable 5 was one of the five agents tested, and the same gap shows up there too. Removing source-code access dropped every model's score, across every metric the benchmark tracks.
That access dependency is the real story here. These agents aren't reasoning their way to a root cause, they're using the codebase to narrow down what to blame. Watch whether oncall-agent pitches start listing full repo access as a requirement in the fine print.
Each link below shares sources, entities, or timing with this story.
Orca supports Copilot / Shared entity: Claude Fable / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Orca supports Copilot); both cover Claude Fable; reported by the same outlet (arxiv.org).
Orca supports Claude Code / Shared entity: Orca / Shared topic / Earlier coverage
Linked by a graph relationship (Orca supports Claude Code); both cover Orca; overlapping topics (agent, best, claude).
Orca supports Claude Code / Shared entity: OpenTelemetry / Shared topic / Earlier coverage
Linked by a graph relationship (Orca supports Claude Code); both cover OpenTelemetry; overlapping topics (access, agent, claude).
Orca supports OpenCode / Shared entity: OpenTelemetry / Shared topic / Earlier coverage
Linked by a graph relationship (Orca supports OpenCode); both cover OpenTelemetry; overlapping topics (agent, claude).
Orca supports Claude Code / Shared entity: Best / Shared topic / Earlier coverage
Linked by a graph relationship (Orca supports Claude Code); both cover Best; overlapping topics (agent, claude).
Orca supports Claude Code / Shared entity: Cohen / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Orca supports Claude Code); both cover Cohen; reported by the same outlet (arxiv.org).
Orca supports OpenCode / Same source domain / Shared topic
Linked by a graph relationship (Orca supports OpenCode); reported by the same outlet (arxiv.org); overlapping topics (agent, claude, task).
Claude Fable built by Anthropic / Same source domain / Shared topic
Linked by a graph relationship (Claude Fable built by Anthropic); reported by the same outlet (arxiv.org); overlapping topics (agent, claude, task).