Opus 4.8 Achieves 0% Uncritical Code Reporting — First Claude Model to Never Silently Pass Flawed Code
VentureBeat·high signal
Opus 4.8 benchmarks show agentic coding 64.3% → 69.2% (beats GPT-5.5's 58.6%), multidisciplinary reasoning 54.7% → 57.9%, and a 10x reduction in overconfidence vs 4.7. It's 4x less likely than Opus 4.7 to let code flaws pass unremarked and scores 0% on uncritically reporting flawed results — the first Claude model to hit that mark. Fast mode is now 3x cheaper at 2.5x speed.