Fetching from the wire…
Agents2026-09-15 · source-backed
Elo-per-token analysis tracks the best solution at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into cross-task ratings, applied to four general agents on four open-ended benchmarks with sessions up to 100M tokens (arXiv 2609.15309). Independent sampling is the reference where Elo grows linearly with log compute. Agents start above it, their marginal gains diminish, and eventually they fall below. The strongest human contestants improve superlinearly over contest time. Practical read: when an agent stalls, extending the budget is the wrong move.
Each link below shares sources, entities, or timing with this story.
157 tasks from 27 repos. Best resolved rate: 29.94%. Agents cause scope creep and regressions by diverging from user intent. Practical implication: agents need explicit scope constraints. arXiv 2509.22237 ---
HarnessOpt-Bench (arXiv 2608.06301) has a frontier LLM act as an optimizer receiving a target agent's seed harness (prompts, tools, control flow, memory, orchestration code) plus graded eval feedback and a fixed evaluation budget, then edits and nominates a candidate scored on...
arXiv 2608.23541 tested 11 verifier-scored optimization tasks under matched compute. Different model families do find structurally different solutions, and then a single round of reading each other's complete outputs erases exactly the diversity that justified using multiple m...
This one annoyed me, in the good way. Researchers took 206 real developer-agent sessions from 13 developers, extracted each developer's preferences from their actual interaction traces via rule-based bootstrapping plus evidence-grounded refinement, then replayed everything aga...
Evaluates how coding agents perform on a language with strict type systems and ownership semantics. Practical for assessing whether your tools are ready for Rust codebases. arXiv 2602.22764 ---
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.