Fetching from the wire…
Research2026-09-03 · source-backed
Trace as State formalizes a mismatch in causal transformers: long-context reasoning often depends on task state discovered only later, and for causal state-update processors, giving the condition first can require exponentially less memory than giving it last. The method reuses collected reasoning traces as a textual proxy for task state and places them before the long-context block on a fresh pass. Against a matched control placing the identical trace after the context, it won 26 of 27 model/task/metric combinations. GraphWalks Parents exact match went 29.2% (initial pass), 43.0% (trace appended), 81.8% (trace prepended) for DeepSeek V4 Pro Preview. This is a prompt-ordering change with no training, and it's the largest free win in this section.
Each link below shares sources, entities, or timing with this story.
State-corruption attacks work because attacker-controlled data makes false claims about the environment that slip past injection filters, since the text reads like an ordinary tool result. PIPES screens each response unit two ways: static field contracts where a schema gives s...
The architecture report describes a 125B sparse MoE activating 6B parameters per token, plus 51B of n-gram embedding tables held off the accelerator in host memory with prefetching (arXiv 2608.30320). Against the prior 397B-A17B model it leads on 8 of 14 pre-training benchmark...
Alibaba International's Accio team open-sourced 107 tasks (53 CLI, 28 browser, 16 file, 10 API/MCP) running against fourteen offline replicas of real business software in a fresh container per task, with verifiers inspecting mock-service state rather than the transcript. Claud...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
arXiv 2607.26935 argues the human-vs-bot label space can't represent agent traffic: an MLP binary classifier misroutes 39.1% of real agent sessions as human, a SAINT transformer 34.5%, while adding an explicit third class yields agent F1 = 1.000 across all 30 runs. Against a f...
Single-shot prompting produced not one valid coverage-producing verification environment on the paper's benchmarks. AgentDV closes the loop with runnability filtering, CSR-grounded checking to cut hallucinated signals, and coverage-guided iteration against measured gaps. Using...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.