Fetching from the wire…
Agents2026-07-28 · source-backed
arXiv 2607.22448 targets on-prem and air-gapped pipelines running quantized 4-8B models behind tool servers, where the dominant failure is omission: an agent reads 20 of 400 records and reports "no anomalies." A 75,476-trial sweep across five models and two engines found a pooled omission rate of 0.62, and 68% originates in deterministic middleware (ingestion, chunking, retrieval plumbing), not model behavior. If you're debugging omissions by swapping models, you're working the wrong 32%. (arXiv 2607.22448)
Each link below shares sources, entities, or timing with this story.
100 real frontier research tasks across seven scientific domains, full lifecycle, 800 annotated trajectories, 45-pattern failure taxonomy (arXiv 2608.14905). The headline isn't a leaderboard, it's a shared deficit: agents can't check what they produced against what they found,...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
arXiv 2608.24306 tests each agent invocation locally for faithfulness against its own inputs, classifying errors as hallucination, uncited input reliance, uncited output or insufficient citation. Applied to three top-ranked open-source deep research systems, nearly every agent...
arXiv 2608.11392 studies what happens when a long-running agent compacts its context: a standing constraint frequently persists as textual residue that no longer governs behavior. Behavioral replay shows models perform the prohibited action far more often with a degraded resid...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
A paper from Stanovsky's group (arXiv:2606.16576) tests whether LLM agents can uncover a hidden deterministic finite automaton through membership and equivalence queries. Performance drops sharply as the automaton grows, and trajectory analysis exposes recurring failures in qu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.