Fetching from the wire…
Agents2026-09-09 · source-backed
It separates risk recognition, pre-action detection and safe task completion instead of collapsing agent safety into one score, generating 1,249 items from 157 sourced tool-use trajectories by constructing paired scenarios differing in whether the request has a safe fulfillment path. Across 20 models, unsafe behavior rises when no safe fulfillment exists, and frontier proprietary models more often recognize the risk and propose alternatives in those cases. arXiv 2609.06783 A benchmark without unsatisfiable requests misses the failure mode that bites in production.
Each link below shares sources, entities, or timing with this story.
FACE-Eval varies where a preference cue is delivered, user message or tool return, across 5,100 samples and 15 open-weight models from 4B to 1.60T parameters (arXiv 2608.29464). Every single model showed lower verbalized commitment for tool-return cues, and unverbalized adopti...
arXiv 2608.11392 studies what happens when a long-running agent compacts its context: a standing constraint frequently persists as textual residue that no longer governs behavior. Behavioral replay shows models perform the prohibited action far more often with a degraded resid...
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18...
Raffi Khatchadourian's replay benchmark measures behavioral instability through three channels that need no access to hidden reasoning text: tool-call trajectories, evidence contacts, decision concentration (arXiv 2607.20491). Across 8,127 replay episodes over 10 models and 3...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
A team at UC Berkeley RDI built an automated scanning agent that achieved near-perfect scores on eight major AI agent benchmarks. SWE-bench Verified: 100%. Terminal-Bench: 100%. WebArena: approximately 100%. FieldWorkArena: 100%. GAIA: roughly 98%. OSWorld: 73%. The agent didn...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.