Skills
A self-evolving runtime beats a direct coding agent 78% vs 59% on SWE-Bench Pro at 1.41x tokens — and gets cheaper as it matures
Argus runs Manager, Planner, Engineer, and Reviewer roles over persistent project state with fixed model weights, self-evolving through runtime state and control policies rather than training. It reports ~78% on SWE-Bench Pro against 59% for Direct Copilot at 1.41x aggregate tokens, 76.8% on AARRI-Bench, and mature waves using 21% fewer solve-input tokens and 15% less active workflow time than startup waves. The reliability machinery is worth copying independently: 34 verifier recoveries and 22 strict review-loop rescues across 254 missions in six real paper pipelines, with 16 stage rollbacks.
↳ Follow the thread