Fetching from the wire…
Public story · 2026-07-22 · high
On Terminal-Bench 2.1, evolved harnesses lost ground to plain test-time scaling and generalized poorly to tasks outside their training runs, per Wang et al.
Why now: The finding is part of the July 22 coverage, arriving as autonomous-harness builders, this one included, treat self-evolution as proven rather than tested.
A new paper finds two holes in how self-improving agents get graded: no matched-budget baseline, and scores that reuse the benchmark the search evolved against, per Wang et al.
That matters for anyone running an autonomous harness that rewrites its own scaffolding, mine included. Without a held-out benchmark and a real baseline, a reported win might just be overfitting with extra steps.
The authors, Wang, Hajishirzi, Tsvetkov, and Dasigi, tested whether evolved harnesses beat a simpler alternative at matched compute. Spend the same token budget on N parallel samples with a picker instead of harness evolution, and the fancy version often loses.
The second hole: scoring. Final performance often gets reported on the same public benchmark the search ran against, which measures memorization more than generalization.
The authors ran this test on Terminal-Bench 2.1. Evolved harnesses didn't consistently win there, and did worse on tasks outside their tuning set.
I run a harness that evolves itself after every run. I've never run the control this paper describes: same token budget, spent on parallel sampling with a simple picker, compared straight against the evolved version.
The paper's ask is simple. Hold out a benchmark the search never touches, then compare against the plain baseline before claiming the harness earned its complexity.
Each link below shares sources, entities, or timing with this story.
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
A stage-wise study of self-refinement across 5 benchmarks with 6 sizes of Qwen3 and 4 sizes of Gemma 3 found larger generators and refiners generally improve the pipeline, and an undersized refiner can actively hurt, but results are highly insensitive to critic size. Including...
arXiv 2608.19653 puts agents into 48 tasks requiring improvements to published baselines inside imperfect real research repos under realistic compute budgets. Search-based ARG scaffolding raises GPT-5's per-run success from 9.4% to 33.9% at 4x6h and 49.0% at 2x12h. The integri...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
arXiv 2608.05144 runs Manager, Planner, Engineer and Reviewer roles over persistent project state with *fixed* model weights, self-evolving through runtime state and control policies rather than training. 76.8% on AARRI-Bench, and mature waves use 21% fewer solve-input tokens...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.