Import AI 468: Intology's Locus Beats the Human Baseline on PostTrainBench+ (51.6% vs 51.1%) Once the 10-Hour GPU Limit Is Removed
Jack Clark's 10 Aug 2026 issue reports Intology's Locus hitting 44.7% on PostTrainBench, then 51.6% on the new PostTrainBench+ variant — which drops the original 10-hour GPU cap — using over 4,000 H100 GPU hours, edging past the 51.1% human baseline. The issue also covers IFP's 23 policy recommendations across seven categories for automated AI R&D (RSI), and a MIT/Columbia game-theory paper, 'Racing to Ruin,' modelling whether competing labs can coordinate a slowdown. The paper's conclusion is the quotable one: 'with low trust, every equilibrium races to ruin' — and transparency alone is ambiguous, since it enables coordination and tempts free-riding depending on the trust level.
↳ Follow the thread
No related signals yet.