A benchmark that blocks chain-of-thought finds Astra chains 34 latent math steps against Sol's 8
LatentMathBench (via r/OpenAI, 95 upvotes)·medium signal
Maarten Baert built LatentMathBench after hearing rumors that GPT-6 Astra uses recurrent depth, specifically to force long reasoning chains in latent space with no chain-of-thought allowed. Astra completed 34 consecutive small-number arithmetic operations; Sol managed 8, and across 24 other models tested the runner-up was Claude Opus 4.6 at 12. The benchmark had been essentially flat until Astra's release, which is the interesting part, though a commenter correctly notes the result does not establish recurrence as the mechanism without testing Huginn or Ouro with recurrency forced to one step.