Recurrent-depth transformers would move reasoning off the token meter, and Jamin Ball thinks that matters more than the architecture
Clouded Judgement's September 4 issue argues looped transformers, where a token passes through the same layers repeatedly, get roughly 2x effective depth without proportional memory growth, decoupling parameter count from compute scaling. The commercially interesting consequence is that reasoning work shifts from visible output tokens into hidden layer passes, which undercuts per-token pricing and reduces chain-of-thought monitorability at the same time. Ball hedges appropriately, calling it possibly 'a modest architectural tweak' analogous to Snowflake separating compute from storage; the same issue puts the cloud software median at 4.4x NTM revenue, 18.1x for high-growth, with 13% median NTM growth and 110% net retention.
↳ Follow the thread