Research
Controlled Study: Instruction Compliance Collapses to Zero at 80 Simultaneous Instructions
A 5-model factorial study (960 calls/model for format, 5,520 calls/model for scaling) finally puts numbers on three prompt-design decisions builders make blind. Perfect-response rate falls to zero by N=80 simultaneous system-prompt instructions, and recall sits near ceiling through 64–128k tokens before degrading sharply — one model lost 48 accuracy points at 128k. Notably, there was no reliable markdown advantage: format winners were model-specific, and fabrication never occurred across 5,760 probes, meaning models drop instructions silently rather than hallucinate. The authors released VeyraBench with the full harness, corpus generator, and raw results.
↳ Follow the thread