Vibe Coding
Opus 5 Agentic-Coding Benchmarks: More Than 2x Opus 4.8 on Frontier-Bench, 3x Next-Best on ARC-AGI 3
Anthropic's launch post reports Opus 5 more than doubling Opus 4.8 on Frontier-Bench v0.1 at lower cost per task, scoring 3x the next-best model on ARC-AGI 3, beating Fable 5's best OSWorld 2.0 result at just over a third of the cost, and landing within 0.5% of Fable 5's peak CursorBench 3.2 score at half the cost. Third-party trackers put absolute Frontier-Bench numbers at 43.3% for Opus 5 at max effort against 18.7% for Opus 4.8, 37.5% for GPT-5.6 Sol and 33.7% for Fable 5, with the widely-shared ARC-AGI 3 figure at 30.2%. Frontier-Bench v0.1 is a 74-task successor to Terminal-Bench 2.1, so this is the terminal-agent benchmark to watch going forward.
Source
↳ Follow the thread