Research
Models Spot Only 9.6% of Implementation Gaps in Research Specs but Fix 80.6% Once You Point Them Out
IdeaAMBIG (arXiv 2609.10539, 2026-09-09) tests whether a research-method specification contains enough information for a coding agent to implement it without inventing assumptions. It holds 660 evidence-grounded instances, 163 real gaps mined from reproducibility reports and GitHub issues plus 497 synthetic gaps injected into codification-ready references. Across 13 LLMs the best model reaches 9.6% Macro Defect Recovery Rate on real instances but 80.6% Macro Clarification Action Success once handed the annotated defect, and an oracle study lifts the downstream codification-ready rate from 14% to 98%. Localization, not repair, is the bottleneck.
↳ Follow the thread