Rauno Arike argues the Astra recurrence panic is overblown: probably 3-4 loops, depth within 2x of GPT-4
In a September 2 LessWrong post, Rauno Arike makes the technical counter-case to the recurrent-depth alarm: the recurrence runs 'along the depth axis rather than across sequence positions,' so it is a looped transformer, not an RNN, and Jakub Pachocki's claim that 'the depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4' bounds it to roughly 3-4 loops. Arike cites Geoffrey Irving's objection that 'to get significant mileage out of bounding the depth of a circuit, you have to bound it very low,' and notes the academic precedent cuts both ways: Huginn trained at 32 loops and scaled to 64 at test time, but the 20B-parameter Loopie uses only 2, hinting at diminishing returns at scale. The open question is whether hundreds of loops ever work.
↳ Follow the thread