Fetching from the wire…
Agents2026-09-03 · source-backed
READY argues an agent can score well and still be undeployable, because the real question is whether it hits a reliability target under acceptable human oversight at tolerable cost. In a clinical-audit study across 16 agent systems and 750 cases, two systems separated by 0.3 points of autonomous accuracy (72.8% against 73.1%) diverged sharply once oversight burden was priced in. The testbed is open and runs on existing agent-evaluation infrastructure.
Each link below shares sources, entities, or timing with this story.
RGA-Designer trains a reward model scoring both task correctness and structural compactness, then fine-tunes a graph generator against it to design communication topologies. arXiv For fan-out agent teams where inter-agent chatter dominates the bill, topology is a cost lever mo...
arXiv 2608.11879 benchmarked Mem0, Hindsight and Mastra Observational Memory across conversations up to 400 turns and 665 LoCoMo questions. Cost models built on conversation length miss badly because internal memory behavior dominates. Break-even against just replaying the ful...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
arXiv 2607.26791 benchmarks post-compromise incident response and reports agents struggle to proactively investigate silent intrusions. They respond to what they're pointed at. Read alongside the July intrusion post-mortem, that argues against putting an agent on the detection...
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
The trajectory-mining pipeline that segments, clusters, and trains a skill-aware policy produced clean skill clusters but only +1.95 points on one benchmark and negligible gains on another (arXiv:2606.20363). A useful negative signal against the hype: generating skills from in...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.