Fetching from the wire…
Agents2026-06-24 · source-backed
PlanBench-XL tests agents on planning under partial observability with mid-task disruptions, conditions closer to real deployment than happy-path benchmarks. NatureBench assembles 90 cross-disciplinary tasks from Nature papers to measure whether coding agents can discover rather than just reproduce known results. Both are author-published and single-source for now, but the framing matters: recovery-from-failure and genuine discovery are sharper bars than the static benchmarks everyone games.
Each link below shares sources, entities, or timing with this story.
Same source / Shared topic
Cite the same source (PlanBench-XL); overlapping topics (agent, benchmark).
Shared entity: PlanBench / Same source domain / Earlier coverage
Both cover PlanBench; reported by the same outlet (huggingface.co); earlier PlanBench coverage from 2026-06-23.
Same source domain / Shared topic / Tension
Reported by the same outlet (huggingface.co); overlapping topics (agent, coding); pushes against this story (competes).
Same source domain / Shared topic
Reported by the same outlet (huggingface.co); overlapping topics (agent, benchmark, condition).
Shared topic
Overlapping topics (agent, benchmark, coding, discovery, evaluation).
Same source domain / Shared topic
Reported by the same outlet (huggingface.co); overlapping topics (agent, benchmark, everyone).
Shared topic / Tension
Overlapping topics (agent, benchmark, coding); pushes against this story (versus).
Overlapping topics (agent, benchmark, coding); pushes against this story (versus).