Fetching from the wire…
Public story · 2026-09-08 · high
Its silent-assumptions score flags agents that guess instead of asking, and InterOPT beat every baseline at recovering them.
Why now: This is a new arXiv entry, and clarification benchmarks for agents are still rare enough that one stands out.
OR-Clarify scores agents on whether they ask for missing details or guess, per the OR-Clarify paper. The benchmark runs each task through a simulated user. It grades slot recovery, when the agent stops asking, whether it made a silent assumption instead, and the cost of the back-and-forth. Most agent benchmarks only check the final answer, so a wrong guess and a right question score the same. This one splits them apart.
The paper's own method, InterOPT, works out which of the withheld details would change the answer, then decides whether to ask again or stop. It beat every baseline on exact slot recovery in the choice-based setting the researchers tested.
The paper doesn't say how InterOPT or the silent-assumptions score hold up outside that choice-based setting, or outside operations-research tasks. Whether it generalizes to a coding agent or a support bot turning a vague request into a spec is untested, not proven.
Each link below shares sources, entities, or timing with this story.
Eight teams per setting formed independently from one base model, each agent keeping a private notebook across ten formation episodes, then role-matched agents were traded between teams (arXiv 2609.05279). Against a placebo reproducing roster-change disruption without changing...
A controlled study ran five Qwen models over eight cases against a DWSIM simulator, 120 slots per arm, with one instruction as the only difference: request a fresh simulation after a substantive modification. No hard gate. Re-verification happened in 94 of 120 guided slots aga...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
arXiv 2608.09902 wraps all 22 boss encounters of Dark Souls: Remastered in a containerized Gymnasium-style benchmark where each step is a real action against the running game. On DSLE-5, an expert system and an evolutionary baseline beat only the tutorial boss (63% and 43% pea...
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.