Preregistered experiment finds structured spec contracts do not beat narrative prose for code generation
A prospective preregistered randomized trial (arXiv:2608.10314, Aug 10) had two LLM snapshots translate five theoretical accounts into code under two presentation formats — structured intervention contracts versus narrative prose — producing 320 programs scored on deterministic metrics. Both primary hypotheses returned NOT_SUPPORTED, with only 19 of 108 criterion evaluations passing, and cross-model identifiability of format sat near chance (AUC 0.469-0.523) against a registered 0.80 threshold. This is a null result, so it bounds rather than overturns the folk practice of rewriting specs into rigid contract format; it suggests the payoff people attribute to structure may come from added content, not formatting. Full design, datasets, and software are on Zenodo.
↳ Follow the thread