Fetching from the wire…
Public story · 2026-07-31 · high
Self-distilling the same model matched externally taught accuracy at 62 times lower cost per query, per CTGT's finance benchmark.
Why now: CTGT published the comparison on July 31, the first cost numbers pitting self-distillation against externally taught distillation in this benchmark.
CTGT trained GPT-OSS-120B on finance-reasoning outputs from the censored DeepSeek V4 Flash. The censorship didn't carry over, per its new research.
That matters for anyone worried that AI labs are laundering bias into open models through distillation. The teacher model showed a 45.45-point censorship gap against its own base version. The distilled student showed no statistically significant gap at all, across 152 matched prompt pairs scored by four independent judges using CTGT's LineageEval framework.
The bigger number is buried further down. CTGT also ran self-distillation, training GPT-OSS-120B on its own outputs instead of the censored teacher's. That version matched the externally taught model's accuracy across all three test seeds. It scored 83.61% on FinanceReasoning at an 8k budget, beating Kimi K3's 81.93% at 62 times lower cost per query.
The censorship result is the reassuring headline. The cost result is the one that's worth acting on. If a self-taught model matches one taught by another lab, that changes the calculus. Paying for that lab's distillation data stops making sense wherever a strong base model already exists. Worth watching whether that holds outside finance reasoning, the only domain CTGT tested here.
Each link below shares sources, entities, or timing with this story.
OpenAI cut GPT-5.6 Luna roughly 80%, from $1 to $0.20 per million input and $6 to $1.20 output. Anthropic priced Opus 5 at $5/$25 per million, half of Fable 5. The trigger is DeepSeek, Zhipu's GLM-5.2 and Moonshot's Kimi K3 landing 60-90% below US flagship pricing, with DoorDa...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.