Fetching from the wire…
Agents2026-08-08 · source-backed
arXiv 2608.06128 attacks confirmation bias in search agents, where the model decides from parametric memory and uses retrieval to confirm itself. CIPO assigns dense turn-level credit specifically to reasoning actions influenced by retrieved information, plus a global outcome reward, with no human process annotations or separate reward model. It reduces prior-driven reasoning across seven in-domain and out-of-domain benchmarks. "It cited a source" and "the source changed its answer" are different measurements, and almost nobody instruments the second one.
Each link below shares sources, entities, or timing with this story.
Raffi Khatchadourian's replay benchmark measures behavioral instability through three channels that need no access to hidden reasoning text: tool-call trajectories, evidence contacts, decision concentration (arXiv 2607.20491). Across 8,127 replay episodes over 10 models and 3...
A June 26 paper (arXiv:2606.26294) describes a self-improving architecture where the agent and the evaluator that scores it evolve together, specifically to avoid the stagnation of optimizing against a fixed, gameable reward. (arXiv) Anyone building a self-improving harness ha...
The paper names it inertia bias: once an agent has produced a query, plan or intermediate conclusion, it judges the consequences of that action less objectively (arXiv 2608.23045). The IBIS benchmark isolates the effect by holding search observations fixed while varying whethe...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
Denser credit assignment attacks the sparse-reward problem that limits RL-trained reasoners (arXiv). If you're training your own reasoning models, segment-level reward is the lever to try when the model gets the right answer through bad steps.
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.