Sources
Prefilling 1% of GPT-5.5 Pro's reasoning into Qwen3.8 raises answer overlap by 18 points, and 27 points on STEM
A v1.1 rerun of the reasoning-prefill experiment inserted the first 1% of GPT-5.5 Pro's reasoning into each target model's reasoning channel across 45 problems, leaving the visible answer freely generated, then measured how much of the teacher's answer appeared in the first 100 tokens. Qwen3.8 A95B jumped from 16.79% to 34.97% source recall (+18.18 pp), rising to +26.99 pp on the 15 STEM problems. DeepSeek V4 Flash moved -1.17 pp and Inkling +0.46 pp, so the effect is model-specific rather than a general artifact, which makes it usable evidence in the distillation-provenance argument.
↳ Follow the thread