Fetching from the wire…
Agents2026-08-10 · source-backed
arXiv 2608.06362 combines variance-reduced outcome estimation with anytime-valid confidence sequences so an eval stops the moment evidence suffices without breaking validity. Across 15 agent configs and 71,439 paired poker hands, AIVAT alone gave 54x median variance reduction, and raw outcomes needed 74x as many hands to hit the same ±1 BB stopping criterion. The reproducibility angle matters more than the cost: it hands a third party everything needed to recheck the verdict at that exact stopping point.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
arXiv 2608.05906 keeps a dual-polarity memory of verified corrections and observed dead ends for Text-to-SQL repair: 66.34% to 69.79% on Spider, 47.35% to 48.44% on BIRD. Then the authors say the quiet part: paired analysis supports the Spider gain but is weak on BIRD, MERIT i...
This method retains four categories of reusable context (task specs, data schemas, tool configs, output constraints) while discarding session-specific reasoning, enabling role-based workspace transfer across users (arXiv:2607.09493). It reports 96% completion versus 79% withou...
APort Vault replays 4,371 human-written attacks from a public CTF against a live payment agent, across 14 models from 8 labs, five policy configurations and two tracks, for 225,964 total evaluations. At Levels 2 through 4, transfers to recipients the passport didn't permit num...
Skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed (arXiv 2609.15396). SkillLift treats ranking as a smoother supervision target than absolute sco...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.