OR-Clarify Tests Whether an Agent Knows It Needs to Ask Before It Starts Modeling
arXiv 2609.05258 introduces OR-Clarify, a benchmark for pre-formulation clarification where each task gives a partial public problem description, withholds structured hidden slots, and evaluates the agent through bounded interaction with a simulated user, measuring slot recovery, stopping behavior, silent assumptions and interaction cost. The companion InterOPT framework identifies unresolved formulation-critical gaps and uses them to decide whether to ask another question or stop, substantially outperforming all baselines on exact slot recovery in the choice-based setting. It sat at 16 upvotes on HuggingFace Daily Papers, and the silent-assumptions metric generalizes well past operations research to any agent that turns vague requests into specs.
↳ Follow the thread