Fetching from the wire…
Research2026-08-08 · source-backed
arXiv 2608.06270 runs a causal audit of visual tool-use, intervening at policy level, trajectory level (corrupting all observations mid-rollout), and step level (counterfactually swapping one observation under a fixed prefix). Across six models and five perception benchmarks: "Calling Without Looking," where returned observations have zero causal effect on the answer, and "Looking Without Planning," where observations are informative but the call schedule is incoherent. Aggregate gains are real but concentrated in a small calibrated minority of rollouts. Most of your crop-and-zoom token spend buys nothing causal.
Each link below shares sources, entities, or timing with this story.
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
A large-scale audit across popular MCP directories found security issues in 5,832 of 9,695 servers, with 2,259 containing exploitable vulnerabilities that go beyond simple auth gaps: arbitrary file access, command injection, SSRF, SQL injection. GBHackers has the writeup. A se...
Most looped-transformer results compare at fixed model size, conflating architecture with extra compute. SMELT matches per-token FLOPs, non-embedding parameters and KV cache against an unlooped baseline, scaling to 54B non-embedding parameters with a separate Chinchilla-style...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
Most visual token pruning runs after the encoder, leaving encoder latency untouched. PACE is training-free in two stages: an Adaptive Pixel Compressor scores visual information density before encoding and downsamples redundant input, then a Dynamic Dual-Attention Extractor kee...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.