ToolGate's Three Acceptance Gates Cut 500 Generated Benchmark Items to 130 Survivors
ToolGate treats every LLM-generated scientific benchmark item as a proposal that must clear three gates: an executable solution script must reproduce the proposed answer when run with the real scientific software, randomized no-tool screening must reject anything a model already solves from the prompt alone, and a tool-using agent must solve each survivor inside a fixed time limit. Instantiated in FEniCSx with 500 generation attempts, local verification retained 478 candidates, two randomized no-tool screens excluded 222 and direct GPT-5.5 API calls at medium reasoning excluded another 121, leaving 135, of which a GPT-5.5 Codex CLI agent with FEniCSx access solved 130. The attrition curve is the useful number for anyone auto-generating evals: roughly three quarters of locally valid items were either trivially answerable or otherwise unusable.
↳ Follow the thread