Agents
Praxist scores 60 MLE-bench medals against a Claude Code baseline's 55, at one twelfth the model spend
Praxist (arXiv 2608.25955, 26 Aug) replaces per-attempt agent memory with a typed evidence graph of findings, lane-structured frontiers and agendas, so later attempts inherit validated mechanisms instead of re-learning them. On the standardized 75-task MLE-bench suite the official grader gave Praxist 60 medals (80.0%), 49 of them gold, versus 55 medals (73.3%) and 34 gold for a Claude Code baseline on Claude Opus 4.8. The recorded spend was $3,054 against $38,370, which is the more useful number for anyone running long autonomous R&D campaigns.
Source
↳ Follow the thread