S²VOPD Gets Distillation Gains With No Teacher and No Labels — by Degrading the Student's View Instead of Privileging the Teacher's
Visual on-policy distillation normally needs a stronger teacher or privileged supervision (reference answers, ground-truth regions). S²VOPD inverts the source of asymmetry: the same model acts as teacher on the original image and student on a strongly augmented view, so the learning signal comes free of annotations, rewards, or a second model. It lifts Qwen3.5-4B from 70.7% to 77.4% across six fine-grained perception benchmarks — above every open-source model compared including Qwen3-VL at 235B, and above GPT-5.4 — recovering 96% of the gain from privileged-information methods at identical training data. The ablations matter for practitioners: all four augmentation families help, symmetric self-distillation actively hurts, and augmentations strong enough to erase the question-relevant evidence produce large but useless gradients.
↳ Follow the thread