Fetching from the wire…
Public story · 2026-08-25 · high
Sonnet 4.6 and GPT-5 barely cared; TypeScript contracts took one weak model's API coverage from 33% to 100%.
Why now: The comparison entered the research corpus on August 25, testing six models still current for agent coding pipelines.
Architecture spec format swung weaker AI models by up to 2.42 points on a quality scale, barely moving frontier models like Sonnet 4.6 and GPT-5. That gap matters most for teams routing agent coding work to cheaper models to cut cost.
The architecture-spec format comparison ran 90 multi-turn agent trials across six models from Anthropic, OpenAI, and Google. It tested five spec formats: informal prose, Mermaid diagrams with ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules.
On Sonnet 4.6 and GPT-5, the five formats scored within 0.17 to 0.92 points of each other. On weaker models the same formats spread 0.83 to 2.42 points, and code-proximate formats recovered most of that gap. TypeScript contracts took the weakest model's API route coverage from 33% to 100%.
Self-validation rates split the same way: 100% on Sonnet, 0% on Gemini Flash. Mid-tier models also burned more tokens than frontier models for worse output when they fell into compilation debugging loops.
Each link below shares sources, entities, or timing with this story.
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.77).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.76).
Same source
Cite the same source (arXiv 2608.21747).
Same source domain / Semantically similar
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.75).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.74).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).
Reported by the same outlet (arxiv.org); covers closely related ground (similarity 0.73).