Fetching from the wire…
Public story · 2026-03-06 · source-backed
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the builder-critical feature: models look up tool definitions on demand instead of loading all schemas upfront, reducing token usage by 47%. This directly addresses the context overhead problem — Claude Code hauls 62,600 characters of tool definitions per turn (55% of context), while Tool Search would let models query tools as needed. Benchmarks: 83.0% GDPval (vs Opus 4.6's 78.0%), 57.7% SWE-Bench Pro, 89.3% BrowseComp (Pro variant). GPT-5.2 deprecated in 3 months. (OpenAI | TechCrunch)
Each link below shares sources, entities, or timing with this story.
Simon Willison uses Claude Code / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (Simon Willison uses Claude Code); both cover Bench Pro, GPT, Opus, OSWorld; reported by the same outlet (openai.com).
Opus built by Anthropic / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Opus built by Anthropic); both cover Bench Pro, Claude Code, GPT, OpenAI; reported by the same outlet (techcrunch.com).
Opus built by Anthropic / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Opus built by Anthropic); both cover Bench Pro, Benchmarks, GPT, Opus; overlapping topics (agent, benchmark, capability, gpt-5, model).
Claude Code uses Opus / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Claude Code uses Opus); both cover BrowseComp, Claude Code, OpenAI, Opus; overlapping topics (benchmark, claude, code, context, model).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Claude Code, Opus, OSWorld, SWE; overlapping topics (benchmark, claude, code, context, token).
Codex competes with Claude Code / Shared entities / Same source / Shared topic / What happened next / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover GPT, Launches, OpenAI, TechCrunch; cite the same source (TechCrunch).
Cursor uses Opus / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (Cursor uses Opus); both cover Bench Pro, GDPVal, GPT, Opus; reported by the same outlet (techcrunch.com).
Claude Haiku benchmarked against BrowseComp / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Claude Haiku benchmarked against BrowseComp); both cover Bench Pro, OpenAI, Opus, SWE; overlapping topics (benchmark, code, model, token).