SimpleOPD Distills Long-Context Reasoning Across Tokenizer Boundaries in Shared Text Space — Intern-S2-Preview Gains 21.2 Points on ProofBench to 55.2, Passing Gemini-2.5-Pro
SimpleOPD (arXiv 2608.14277, submitted Aug 14) tackles the practical blocker in on-policy distillation: teacher and student rarely share a tokenizer. It performs distillation in shared text space to sidestep the mismatch entirely, then adds a student reference KL loss and masks the advantages of special termination tokens to stop runaway generation length. Intern-S2-Preview reached 55.2 on ProofBench, a 21.2-point improvement that puts it past Gemini-2.5-Pro, with gains replicated across Qwen3, Qwen3.5, Intern-S2, GLM-4.7 and Gemma-4 and generalizing to HLE and HiPhO. Tokenizer-agnostic matters because it means you can distill from whichever teacher is strongest rather than whichever one shares your vocabulary.
↳ Follow the thread