Fetching from the wire…
Agents2026-06-22 · source-backed
arXiv 2606.18847 introduces a benchmark and models for agents that must carry state across many steps instead of acting one-shot, targeting the durability-of-state gap most agent evals skip. It's aimed at realistic, hours-long tasks rather than single-turn actions. If you're building agents that run for hours and quietly lose the plot halfway through, this is the eval that measures the failure you're feeling.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, eval); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (action, agent, benchmark); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, embodied, failure); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (action, agent, failure); pushes against this story (but).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (action, agent, benchmark, failure).
Reported by the same outlet (arxiv.org); overlapping topics (acting, action, agent, benchmark).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, eval); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, failure); pushes against this story (against).