Research
Agents Convert Tokens Into Progress Faster Than Independent Sampling at First, Then Fall Below It
Elo-per-token analysis tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into cross-task Elo ratings, applied to four general-purpose agents on four open-ended benchmarks with sessions up to 100M tokens. Independent sampling gives a theoretically characterized reference where Elo grows linearly with log compute. Agents initially beat that reference but their marginal gains diminish and eventually drop below it, while the strongest historical human contestants improve superlinearly over contest time, which argues against simply extending an agent's budget when it stalls.
↳ Follow the thread