Raschka pushes back on the 'hidden reasoning' reading of GPT-6 Astra and puts numbers on looped transformers
Raschka argues Astra's gains most likely come from training recipe and data rather than architecture, noting OpenAI's chief scientist saying compute depth stays within a factor of two of GPT-4. He walks through weight-reused looped stacks going back to Universal Transformers (2018), with Nanbeige4.2-3B applying 22 blocks twice and Ouro-Thinking 2.6B applying 48 blocks four times, plus Mixture-of-Recursions for per-token loop counts. On the claim that looping hides reasoning traces, he counters that shorter chains reflect capability and post-training tuning, citing GPT-5.6 Sol using 80% more tokens than Luna at similar performance, and cites SMELT (2026) measuring 6.8 to 18% less training compute for equal validation loss.
↳ Follow the thread