Fetching from the wire…
Agents2026-09-25 · source-backed
arXiv 2609.28585 names the mechanism "persistent billable state," a tool output the host runtime carries forward into later turns where the provider meters it again. Across 243 executions on six model families, DOW-BENCH saw cumulative session input reach 14,293x the first call's input, and keeping raw history raised mean session cost 21.2 to 35.9%. The defense is host-side, not prompt-side: deterministic history compression (10 to 11 of 12 history-dependent tasks still succeeded, against 2 of 12 under plain deletion) plus four invariants bounding prompt mass, context growth, recursion and cumulative spend. Those four contained every recurring attack in a 123-run replay corpus.
Each link below shares sources, entities, or timing with this story.
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
ByteDance Seed's GST-Bench covers 6,790 minutes of synthetic video with human-verified questions, isolating a specific failure: models handle local spatial relations competently but can't consolidate long-horizon observations into a globally consistent scene. The ~36-point gap...
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
Unreal Labs open-sourced it under MIT in Go. Each tool call is recorded as in-progress in the session log while the tool runs in the background, so the model never spends tokens waiting or polling. Same 57.9% pass rate as Codex at $1,428 against $2,350, and 65.8% on SWE-Atlas...
arXiv 2609.20152 targets a measurement gap: end-to-end voice benchmarks mix recognition and model errors into one number, while LLM benchmarks isolate the model and drop the conditions that make phone calls hard (arXiv). The benchmark runs the model with transcription errors p...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.