Fetching from the wire…
Public story · 2026-08-24 · high
The paper's seven-part framework says current benchmarks track final outcomes and miss recovery and governance almost entirely.
Why now: Agent evaluation is still catching up to how differently these systems are built, and this survey is an early attempt to name what a fair comparison requires.
A new survey called Terminal Agents argues that agent behavior depends more on harness and runtime than on the model underneath it.
That reframing matters for anyone comparing agent products or benchmarks. A shared task score can hide very different setups, from a bare model with a shell to a fully tooled agent with retries and sandboxing.
The paper defines terminal agents as systems whose dominant loop runs through command execution and stateful interaction with an environment. These are the agents used for coding and operations work. It sorts agent competence into seven dimensions instead of one pass-or-fail score. Current evaluations reward whether an agent finishes the task, but measure process quality, error recovery and governance far less consistently.
Its recommendation is procedural, not statistical. Report the system and runtime conditions next to the score, and keep replayable traces of the run. The paper backs this with matched system comparisons that showed performance swinging by benchmark. That limited how much of a result could be attributed to any single component.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, alone, behavior); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, behavior, condition); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, arguing); traces where this leads (which means).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (behavior, system); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, comparison); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (agent, behavior); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (agent, component); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, alone); pushes against this story (but).