Sources
AI4AI at Test-Time: A Strong Model Rewriting a Weak Model's Harness Nearly Doubles Theory-of-Mind Accuracy (0.49 → 0.91) With No Retraining
Cheng Qian, Heng Ji, Silvio Savarese and co-authors (arXiv 2608.12307, submitted August 12) show that capability transfers from strong to weak models through the inference harness rather than through parameters. Gains came from offloading unstable reasoning into deterministic code, benchmark-specific routing, and strict answer-format enforcement — explicitly not from making the target model reason longer or sample more. Weaker models saw the largest gains, which makes harness engineering a direct substitute for distillation when you cannot retrain the model you are stuck with.
↳ Follow the thread