← The Wire
Source trail

Essa Mamdani (VulcanBench coverage)

Public MindPattern findings, entities, and graph evidence that cite this source.

Findings
1
All-time hits
1
High value
0
Last seen
2026-08-03

Related findings

  1. 2026-08-03 / AGENTSVulcanBench Suite 3: DeepSeek V4-Flash ties Grok 4.5 at 91% pass@1 for a third the cost, and high effort makes it worseMorgan Linton released Eval Suite 3 of VulcanBench on August 1 — 23 frontier-hard tasks drawn from real merged open-source PRs, Docker-sandboxed, pass@1 with no retries or majority voting, each run annotated with actual dollar cost. Three entries tie at 91%: DeepSeek V4-Flash at medium effort ($2.04), Grok 4.5 at medium ($6.67) and Grok 4.5 at high ($13.16). Two counterintuitive patterns: Grok 4.5 is flat between medium and high effort, and DeepSeek V4-Flash actually drops from 91% at medium to 87% at high, so more reasoning tokens can degrade engineering decisions. Treat as one snapshot from a single evaluator — Kimi K3 and Claude Haiku 4.5 ran with partial task coverage — but the effort-tuning finding is worth testing on your own harness.
Open latest cited source