Fetching from the wire…
Agents2026-09-15 · source-backed
Commitment-Frontier Residual Completion formulates model-to-model handoff as commitment-constrained residual completion: freeze a residual contract from accepted progress, close the successor's continuation into an evidence-linked graph, admit execution only when the remainder is covered (arXiv 2609.13800). Across five environments and two same-provider model pairs it matches strong full-task agents on macro accuracy at 22.0% to 34.6% of inference cost, and cross-provider transfer holds. Three enforced invariants: target-before-proposal, whole-proposal-before-authority, live-evidence-before-success.
Each link below shares sources, entities, or timing with this story.
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
RGA-Designer trains a reward model scoring both task correctness and structural compactness, then fine-tunes a graph generator against it to design communication topologies. arXiv For fan-out agent teams where inter-agent chatter dominates the bill, topology is a cost lever mo...
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
This method retains four categories of reusable context (task specs, data schemas, tool configs, output constraints) while discarding session-specific reasoning, enabling role-based workspace transfer across users (arXiv:2607.09493). It reports 96% completion versus 79% withou...
I've spent the better part of a year building typed tool definitions. JSON schemas, argument validation, careful descriptions so the model picks the right one. Everyone I know building agents has done the same thing. MCP made it a standard. A new controlled study says we may h...
In an Oracle-to-PostgreSQL migration study over 1,006 PL/SQL files, cross-agent experiments ran 1,802 Oracle scripts through Amazon Kiro, Gemini and Copilot. The worst replicated case was Gemini consuming a Kiro-origin specification directly: Token F1 of 0.035, SQL validity 2....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.