Architecture Spec Format Barely Matters for Sonnet 4.6 and GPT-5, But Swings Weaker Models by up to 2.42 Points
A controlled experiment ran 90 multi-turn agent trials across six models from Anthropic, OpenAI, and Google, comparing five informationally equivalent architecture specification formats (informal prose, Mermaid plus ADRs, OpenAPI, C4/Structurizr DSL, and TypeScript interface contracts with ArchUnit-style rules). On the strongest models the quality spread across formats was 0.17-0.92 points, but on weaker models it reached 0.83-2.42, with code-proximate formats recovering most of the capability gap; TypeScript contracts took the weakest model's API route coverage from 33% to 100%. Self-validation rates collapsed from 100% on Sonnet to 0% on Gemini Flash, and mid-tier models burned more tokens than frontier models for worse output when they fell into compilation debugging loops.
↳ Follow the thread