Hacker News
Netlify Ran One Prompt Through 11 Models: Claude Opus 5 Burned 519 Credits, DeepSeek V4 Flash Did It for 2.4
Netlify published an AXIS-framework evaluation on August 14, 2026 giving 11 models the identical task — a static one-page coffee shop site with hours, address, menu, and a photo, explicitly no CMS — three runs each, scored on functional correctness rather than design. Average credit cost spanned more than 200x: DeepSeek V4 Flash 0731 at 2.4, Kimi K2.7 Code at 19, GLM 5.2 at 27, GPT 5.6 Terra at 39, Gemini 3.1 Pro at 53, Kimi K3 at 102, GPT 5.6 Sol at 141, Claude Sonnet 5 at 143, and Claude Opus 5 last at 519. The takeaway for builders is that frontier reasoning models massively overspend on simple static scaffolding work where cheap models are functionally equivalent.
↳ Follow the thread