Fetching from the wire…
Public story · 2026-08-25 · high
Frontier models barely budged across formats, but typed interface specs took one weak model's API coverage to 100%, up from a 33% baseline.
Why now: This test of six model tiers is new to the August 25 coverage of AI coding research.
Six AI coding models ran 90 trials testing whether spec format changes what they build, per a new spec-format study. Format barely moved the strongest models, whose scores spread just 0.17 to 0.92 points across five formats. Weaker models swung by up to 2.42 points, wide enough to change how much of a spec gets implemented.
The team tested five formats carrying the same information: prose, Mermaid diagrams with decision records, OpenAPI specs, C4/Structurizr DSL, and TypeScript interfaces with ArchUnit-style rules. Code-proximate formats closed most of the gap on weaker models.
The starkest number came from API route coverage. The weakest model covered just 33% of required routes from prose specs and 100% from typed TypeScript interfaces.
Self-checking behavior split too. Sonnet caught its own spec violations 100% of the time; Gemini Flash caught none. Mid-tier models burned more tokens than the frontier models and still produced worse output. They got stuck rewriting code until it compiled, without fixing the architecture violation underneath.
Each link below shares sources, entities, or timing with this story.
Claude Code uses Sonnet / Shared entity: OpenAPI / Earlier coverage
Linked by a graph relationship (Claude Code uses Sonnet); both cover OpenAPI; earlier OpenAPI coverage from 2026-07-30.
Claude Code uses Sonnet / Shared entity: Sonnet / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Sonnet); both cover Sonnet; reported by the same outlet (arxiv.org).
Claude Code uses Sonnet / Shared entities / Earlier coverage
Linked by a graph relationship (Claude Code uses Sonnet); both cover Self, TypeScript; earlier Self coverage from 2026-08-21.
Linked by a graph relationship (Claude Code uses Sonnet); both cover Sonnet, TypeScript; earlier Sonnet coverage from 2026-07-19.
Anthropic released Sonnet / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Sonnet); both cover Self, Sonnet; overlapping topics (architecture, model).
Claude Code uses Sonnet / Shared entity: Sonnet / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Sonnet); both cover Sonnet; overlapping topics (model, point).
Claude Code uses Sonnet / Shared topic
Linked by a graph relationship (Claude Code uses Sonnet); overlapping topics (architecture, model).
Claude Code uses Sonnet / Shared entity: TypeScript / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Sonnet); both cover TypeScript; earlier TypeScript coverage from 2026-04-16.