Fetching from the wire…
Public story · 2026-07-14 · high
Grok Build hit 83.3% on Terminal-Bench 2.1 using about a quarter of the output tokens Opus 4.8 needs for the same coding task.
Why now: Grok 4.5 launched July 8, making this the first direct price comparison against Claude's Opus 4.8.
xAI priced a Grok Build coding task at $2.49 on July 8, against $11.80 for Claude Code on Opus 4.8, per TechTimes. That price gap is what budget-conscious teams will care about most. Grok Build runs at roughly a fifth of Claude Code's cost per task, on scores close to frontier level.
Grok Build hit 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro. It got there on about 15,954 output tokens, a quarter of the roughly 67,020 tokens Opus 4.8 used on the same task.
The catch, flagged separately from the benchmark scores: hallucination rates are up. A cheap task that comes back as a confident wrong patch you merge anyway isn't cheap at all.
There's a second bill. The Grok Build CLI was caught silently uploading entire repos, Git history included, to an xAI cloud bucket. The $2.49 sticker price doesn't include what it costs when your codebase leaves the building without you choosing to send it.
The move isn't switching wholesale. Run Grok on work where a hallucination is cheap to catch, refactors with solid test coverage, mechanical migrations, throwaway prototypes.
Keep the high-stakes, hard-to-verify work on the model you trust. My bet: the repo-upload story costs Grok Build more adoption than the hallucination rate does.
Trust in data handling is harder to win back than trust in a single patch. Watch whether xAI changes the CLI's default upload behavior, or teams route around it instead.
Each link below shares sources, entities, or timing with this story.
Grok Build competes with Cursor / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Grok Build competes with Cursor); both cover Bench, Bench Pro, Cursor, Grok; overlapping topics (benchmark, cost, grok, model, opus).
Grok Build competes with Cursor / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Grok Build competes with Cursor); both cover Bench Pro, Cursor, Grok, July; overlapping topics (benchmark, decision, grok, opus, price).
Grok Build competes with Claude Code / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Grok Build competes with Claude Code); both cover Bench, Claude Code, July, Opus; overlapping topics (benchmark, cost, model, opus, price).
Linked by a graph relationship (Grok Build competes with Claude Code); both cover Bench, Claude Code, Grok, Grok Build; overlapping topics (cost, grok).
Linked by a graph relationship (Grok Build competes with Claude Code); both cover Bench, Claude Code, Opus, Security; overlapping topics (benchmark, model, task, token).
OpenAI uses Claude Code / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI uses Claude Code); both cover Bench, Opus, SWE, Terminal; overlapping topics (cost, model, opus, task, token).
Grok Build competes with Claude Code / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Grok Build competes with Claude Code); both cover Claude Code, July, Opus, SWE; overlapping topics (benchmark, model, opus).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Bench, Bench Pro, Opus, SWE; overlapping topics (benchmark, model).