Fetching from the wire…
Research2026-09-19 · source-backed
arXiv 2609.02852 introduces "linguistic illegibility," the gap between what a model says and the math it performs over activation spaces, and argues chain-of-thought monitoring, constitutional self-critique and activation probing are unsound in principle as security mechanisms. The proposed alternative is sandboxing resting on techniques independent of what the model reports, specifically taint tracking over which system states a model's output influenced. arXiv Spend the effort on the boundary, not on reading traces.
Each link below shares sources, entities, or timing with this story.
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
A stage-wise study of self-refinement across 5 benchmarks with 6 sizes of Qwen3 and 4 sizes of Gemma 3 found larger generators and refiners generally improve the pipeline, and an undersized refiner can actively hurt, but results are highly insensitive to critic size. Including...
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
Reasoning models engage in "performative" CoT — the model's final answer is decodable from activations far earlier than visible CoT suggests. Activation probing enables up to 80% token reduction on MMLU. Critical for safety monitoring: visible reasoning may not reflect actual...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.