Self-Study Reconsidered exposes the hidden fragility of training on self-generated QA
arXiv·medium signal
Teaching models from synthetic question-answer pairs a model generates about its own documents is treated as neutral preprocessing, but this work shows the generation step is an implicit policy that both selects which evidence becomes training signal and decides how it's answered — and is fragile at both stages. This matters for anyone doing self-distillation, knowledge compression, or synthetic-data fine-tuning. The takeaway: your QA-generation prompt silently determines what your student model learns.