Fetching from the wire…
Public story · 2026-08-17 · high
ScienceFlow beats the prior best MLE-bench score by 4.92 points, letting agents re-anchor on saved checkpoints instead of retrying from scratch.
Why now: ScienceFlow's paper, arXiv 2608.14354, posted in August 2026 with a new best score on MLE-bench's hardest setting.
ScienceFlow hit 70.22% Any-Medal on the full MLE-bench inside a 24-hour budget, per the arXiv paper describing it, written by a team of 19 authors. MLE-bench measures whether an agent can carry a research problem to a medal-worthy result inside a fixed time budget. A 4.92-point jump on its hardest setting is the difference between an agent that recovers from a stalled direction and one that just starts over.
ScienceFlow breaks a research run into research segments, each with a recoverable, executable state. ESTRA, short for Executable-State Transition through Re-Anchoring, decides what happens when a segment stalls. It can keep pushing on the current direction, or jump back to an active or archived checkpoint and continue from there.
The paper doesn't say how much of the 4.92-point gain comes from ESTRA's re-anchoring decision versus the segmenting alone. That gap matters for anyone trying to copy the approach. It also doesn't report what keeping all those recoverable states costs in compute against a simpler retry loop.
The bet worth testing: pull re-anchoring out of a comparable system, run retries only, and check whether the score falls below 70.22%. Watch whether other agent benchmarks start reporting results by segment-recovery mechanism instead of one lumped number.
Each link below shares sources, entities, or timing with this story.
Shared entity: MLE / Same source domain / Shared topic / Earlier coverage
Both cover MLE; reported by the same outlet (arxiv.org); overlapping topics (agent, budget).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, author, budget, state); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (active, agent, beat, state); pushes against this story (vs).
Reported by the same outlet (arxiv.org); overlapping topics (author, beat, benchmark, budget); pushes against this story (against).
Shared entity: Executable / Same source domain / Earlier coverage / Tension
Both cover Executable; reported by the same outlet (arxiv.org); earlier Executable coverage from 2026-08-03.
General Intuition partners with Medal / Shared entity: Medal / Earlier coverage
Linked by a graph relationship (General Intuition partners with Medal); both cover Medal; earlier Medal coverage from 2026-06-26.
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, state); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, beat, benchmark); pushes against this story (against).