Research
ASI-Bench Withdraws Human Guidance Mid-Project and Agent Scores Fall From 50.91 to 26.62
ASI-Bench (arXiv 2608.17271, 2026-08-18) took 40+ experts and 31,000+ human hours to build 60 project-level research tasks across 11 scientific domains, and its design choice is the interesting part: it progressively removes methodological guidance within the same project to measure how far a system gets on its own. Across 18 agent-model configurations the average score drops from 50.91 with full methodological guidance, to 29.10 when only the method is specified, to 26.62 when the agent must choose the method itself. Tasks pass expert review, AI-assisted auditing, sandbox execution and scorer validation, and submissions are open at asibench.apexin.ai.
↳ Follow the thread