Fetching from the wire…
Public story · 2026-07-31 · high
ORCA-bench found the weakest of five frontier agents invented oncall root causes 40% of the time, using a live incident system with full source access.
Why now: The benchmark's numbers landed in the July 31 briefing, as teams weigh putting agents on oncall rotations.
The best of five frontier agents solved just 25.3% of realistic root-cause tasks on a new oncall benchmark, per ORCA-bench's authors. That's the number SRE and platform teams should plan around before handing an agent the pager. On the Hard tier, the top score drops to 10%. The weakest model invents an implausible root cause in 4 of every 10 reports.
ORCA-bench runs on a live, OpenTelemetry-instrumented microservice system. It provides six days of real metrics, logs and traces from Prometheus, Jaeger and OpenSearch, plus full access to the source code. Agents work through 1,079 tasks that vary how specific the incident report is and whether more than one fault is firing at once. Ground truth came from expert SREs, with agreement scored at a Cohen's weighted kappa of 0.90, so the scoring isn't the weak point.
Claude Fable 5 was one of the five agents tested, and the same gap shows up there too. Removing source-code access dropped every model's score, across every metric the benchmark tracks.
That access dependency is the real story here. These agents aren't reasoning their way to a root cause, they're using the codebase to narrow down what to blame. Watch whether oncall-agent pitches start listing full repo access as a requirement in the fine print.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
+7,546 stars this week for an "Application Development Environment" running each concurrent agent in its own worktree. Works with any terminal CLI agent — Claude Code, Codex, OpenCode, Pi, Cursor, Copilot, Grok, 30+ others — across macOS, Windows, Linux, iOS via App Store/Test...
Amazon CloudWatch launched a purpose-built view ingesting OpenTelemetry metrics from coding agents: total tokens consumed, cost, active users, sessions, cache hit rate, active hours, sitting alongside your existing operational data. Claude Code telemetry collects with no extra...
This one annoyed me, because I've been running the losing pattern. SWE-QA (arXiv 2608.01507) compares the sub-agent grep pattern that Claude Code, Codex and Antigravity all ship by default against a pre-built semantic index over the same repository. Semantic search answered 65...
I check Product Hunt maybe once a week and usually regret it. Today's board is worth reading as market structure. The July 30 leaderboard: SKI at 277 upvotes (free voice input for Claude Code and Codex). AI Search Console at 249 (prompt analytics and citation mapping). Memmy A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.