Distilling a Censored Chinese Model Into GPT-OSS Did Not Transfer the Censorship
CTGT·high signal
CTGT published research on July 29 testing whether political censorship survives distillation, training GPT-OSS-120B on DeepSeek V4 Flash outputs in a finance-reasoning domain. Their LineageEval — 304 prompts as 152 matched pairs, scored by four independent judges — measured a +45.45 censorship gap in the teacher but no statistically significant behavioral difference between the distilled student and the untouched base. The larger builder takeaway is buried in the numbers: self-distillation matched externally-taught performance across all three seeds, hitting 83.61% on FinanceReasoning at 8k budget and beating Kimi K3's 81.93% at 62x lower cost per query.