Fetching from the wire…
Agents2026-09-07 · source-backed
Across 48 synthetic histories and two sub-10B open-weight models, a fixed-schema knowledge graph moved by +0.0004 ± 0.0020 accuracy through a writer swap, while compressed natural-language notes shifted +9.91 or -13.28 points depending on migration direction (arXiv 2609.05339). A 50/50 mixed embedding index captured only 4.96 of the 11.90-point gain available from full re-embedding, so partial re-embedding is close to worthless. Store-only repair of notes never reached 90% recovery in any of the 48 cases; keeping the raw source history recovered 34 of 48. Three rules fall out: normalize to a schema, never partially re-embed, keep the transcripts you compressed from.
Each link below shares sources, entities, or timing with this story.
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
SIR composes stealthy OS-level injections from a small plain-language library of reusable principles, then diagnoses failed trajectories to distill new named strategies and reapply them (arXiv 2608.30207). Scored with a deterministic oracle checking filesystem, service and per...
Fixed synthesis recipes apply the same prompting policy to every seed regardless of whether the current policy needs harder or easier tasks (arXiv 2608.14312). Envs-FORGE estimates per-seed pass rates, scores six projection-direction actions around a target learning frontier,...
Store reusable procedures plus bindings, applicability conditions, and verification requirements instead of the raw trace. That gained 10.7 points of success on WebArena, WorkArena, and AppWorld while cutting online tokens 48.9% (arXiv). Direct trace reuse gets worse as your t...
Holding retrieval, target state, model, decoding and tool budget fixed, researchers compared how a retrieved memory gets used. A target-bound note recording a reusable procedure, bindings to recover, applicability conditions and verification requirements hit 62.3% average succ...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.