Fetching from the wire…
Public story · 2026-07-31 · high
Self-distilling the same model matched externally taught accuracy at 62 times lower cost per query, per CTGT's finance benchmark.
Why now: CTGT published the comparison on July 31, the first cost numbers pitting self-distillation against externally taught distillation in this benchmark.
CTGT trained GPT-OSS-120B on finance-reasoning outputs from the censored DeepSeek V4 Flash. The censorship didn't carry over, per its new research.
That matters for anyone worried that AI labs are laundering bias into open models through distillation. The teacher model showed a 45.45-point censorship gap against its own base version. The distilled student showed no statistically significant gap at all, across 152 matched prompt pairs scored by four independent judges using CTGT's LineageEval framework.
The bigger number is buried further down. CTGT also ran self-distillation, training GPT-OSS-120B on its own outputs instead of the censored teacher's. That version matched the externally taught model's accuracy across all three test seeds. It scored 83.61% on FinanceReasoning at an 8k budget, beating Kimi K3's 81.93% at 62 times lower cost per query.
The censorship result is the reassuring headline. The cost result is the one that's worth acting on. If a self-taught model matches one taught by another lab, that changes the calculus. Paying for that lab's distillation data stops making sense wherever a strong base model already exists. Worth watching whether that holds outside finance reasoning, the only domain CTGT tested here.
Each link below shares sources, entities, or timing with this story.
Kimi K3 competes with OpenAI / Shared entities / Earlier coverage
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Chinese, GPT, Kimi K3; earlier Chinese coverage from 2026-07-21.
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover Chinese, GPT, Kimi K3; earlier Chinese coverage from 2026-07-20.
DeepSeek released deepseek-v4-flash / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover Chinese, GPT; overlapping topics (chinese, deepseek).
DeepSeek released deepseek-v4-flash / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover Chinese, GPT; earlier Chinese coverage from 2026-04-20.
DeepSeek released deepseek-v4-flash / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover GPT; overlapping topics (cost, deepseek).
Kimi K3 competes with DeepSeek V4 / Shared entity: Kimi K3 / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 competes with DeepSeek V4); both cover Kimi K3; overlapping topics (budget, cost, deepseek).
DeepSeek released deepseek-v4-flash / Shared entities / Earlier coverage
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover GPT, Kimi K3; earlier GPT coverage from 2026-07-20.
DeepSeek released deepseek-v4-flash / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover GPT; overlapping topics (budget, cost, deepseek).