Repeating only the procedural instruction lifts compliance from 90.22% to 93.17% while leaving accuracy untouched at 60.21%
arXiv 2609.04024 tests instruction duplication across seven instruction-tuned models and 16,800 scheduled generations, and finds that going from one copy to two raises a deterministic eight-test compliance diagnostic by 2.95 points, eliminating 30.2% of the residual failures, while final-answer accuracy stays exactly 60.21%. The effect is placement-sensitive: in a downstream repair loop that consumes the exposed trajectory, a trailing duplicate moved one endpoint from 84.2% to 97.1% but cut another from 78.6% to 73.8%. This separates cleanly from the older Google prompt-repetition result (arXiv 2512.14982), which claimed accuracy gains; here the win is procedural fidelity, which is what matters when an agent's steps are parsed or repaired by another system.
↳ Follow the thread