Fetching from the wire…
Research2026-07-28 · source-backed
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: preserve local reward with no team success, or query but fail to communicate anything that changes the final decision. Applying SFT, RLOO, sequential SFT+RLOO and GEPA as diagnostics produced heterogeneous effects, with GEPA sometimes raising team success while eliminating the costly queries entirely. Optimization can move aggregate reward without recovering the mechanism. Direct warning against scoring multi-agent systems on outcome alone. (arXiv 2607.23982)
Each link below shares sources, entities, or timing with this story.
NPO iteratively revises a prompt using a teacher model and rollout feedback, with no elaborate search (arXiv 2608.27266). It matches or exceeds GEPA at lower rollout cost, and its advantage widens with stronger teachers, which suggests teacher reasoning substitutes for optimiz...
Multiple independent models train against each other with peer-derived rewards and no ground-truth labels, gaining 3.0-8.6% across seven text benchmarks and 2.3-7.2% across four multimodal ones. The mechanism claim matters more than the numbers: varying architectures, model si...
Flagged by The Batch #365, arXiv 2605.08382 measures the benign case, not adversarial red-teaming. Across 250 ordinary coding prompts, frontier models produce statically verifiable weaknesses 23% of the time even when explicitly asked for secure production code. 12.7% of outpu...
Mohamed Jouini evaluates seven agentic strategies on IaC-Eval v2, 186 AWS/Terraform tasks with Rego v1 intent policies (arXiv 2607.20478). ReAct with MCP or ChromaDB-backed RAG lifts Qwen2.5-Coder 7B from 14.0% to 45.7%; iterative refinement on verifier feedback reaches 62.9%...
Released August 3, experimental dspy.Flex moves program structure into the optimizer's search space: given a signature, GEPA rewrites predictors, control flow, and the Python/LM call balance against your metric, with optimizer-authored source always running inside a CodeInterp...
LLMs are brittle to renamed nodes and reworded formulations in graph reasoning, and the standard fix is throwing a multi-agent system at the parsing failures. GRAIN is a single RL-trained agent modeling reasoning as semantic parsing plus tool execution, rewarded by a Structure...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.