Agents
AutoSciRub writes the grading rubric before the research runs and gains 16.8 points on AstaBench discovery tasks
Open-ended research instructions rarely specify which analyses or success criteria matter, so agents skip analyses, pick wrong methods, or overclaim. AutoSciRub decomposes the instruction into atomic scientific goals, grounds them in literature and task-visible data, and synthesizes verifiable criteria into an executable rubric used to guide execution and drive targeted revision. On ResearchClawBench it gains 2.08 points averaged across three backbone LLMs under a fixed Codex harness and 2.95 across three harnesses on a fixed backbone; on a 20-task AstaBench E2E Discovery subset it gains 16.8 points across three harnesses without reducing task completion. Code is at zjunlp/AutoSciRub.
Source
↳ Follow the thread