Fetching from the wire…
Agents2026-09-18 · source-backed
arXiv 2609.20045 audits context compression with paired histories that share the same current answer, receive the same future update, then require different answers. A deterministic frontier selector scored 96/96 strict reveal accuracy on DeepSeek but 82/96 on GLM, a structured writer managed 56/96, and a record-level audit found 26 and 25 memories that were well formed and semantically wrong. Identifier renaming dropped frontier late-reference adequacy from 8/8 to 94/320 transformed instances. If you compact long agent sessions, the failure won't show up in the turn where compaction happens.
Each link below shares sources, entities, or timing with this story.
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
Numbers first, because they're the whole argument. On AppWorld's 168 tasks with DeepSeek-V3.2: ALTK-Evolve hit 89.3% goal completion at 263K tokens per task. ACE hit 80.4% at 634K (IBM Research on HF). Nine points better, 41% of the tokens. On the weaker gpt-oss-120b: 56.0% at...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
The attack frames protected attributes as operationally required, so the model includes them in otherwise valid tool-call arguments. Across six pressure levels, four privacy-policy levels and five DeepSeek and Claude configurations over 120 calls, stronger privacy instructions...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.