Fetching from the wire…
Agents2026-07-29 · source-backed
arXiv 2607.25825 targets the layer builders actually control, arguing today's harnesses use hand-crafted or globally fixed policies that mismatch task demands and burn compute. It estimates workflow advantage from confidence-weighted execution evidence and only applies adjustments backed by sufficient expected advantage. Across long-horizon information-seeking, software engineering, and terminal tasks, it preserves task success while substantially cutting tokens and wall time.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26598 targets the failure where an agent recovers from an error within an episode but hits the identical failure in later tasks, because post-episode feedback never revises the persistent harness. Guided by a domain-level Evolution-SOP, it writes episodic memory rec...
HarnessOpt-Bench (arXiv 2608.06301) has a frontier LLM act as an optimizer receiving a target agent's seed harness (prompts, tools, control flow, memory, orchestration code) plus graded eval feedback and a fixed evaluation budget, then edits and nominates a candidate scored on...
It treats the executable runtime, context construction, tool mediation, action validation, execution recovery, as the thing to learn. A separate harness engineer converts batches of target-agent failures into validated executable patches, with same-batch reruns of the frozen t...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
ARCHER is a test-driven multi-agent program-synthesis harness for building-compliance checking. Evaluating six harnesses of increasing agentic sophistication across four backbone models, deterministic orchestration won for *every* backbone, improving mean union accuracy 82% ov...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.