Fetching from the wire…
Research2026-09-24 · source-backed
Patching the residual stream at each stated intermediate step with activations from a counterfactual run, on 2-6 hop lookup tasks. The standard behavioral text-editing test reports 88.2% on the same items (arXiv). Qwen3-1.7B drops to 54.8%. CoT monitoring on small models rests on weaker ground than the behavioral metrics suggest, and the behavioral metrics are what most monitoring systems use.
Each link below shares sources, entities, or timing with this story.
Reasoning models engage in "performative" CoT — the model's final answer is decodable from activations far earlier than visible CoT suggests. Activation probing enables up to 80% token reduction on MMLU. Critical for safety monitoring: visible reasoning may not reflect actual...
Activation probes are usually evaluated against agents who don't know they're monitored, which is a generous assumption. This study held models, probes and thresholds fixed and varied only the disclosure: nothing, monitor present, or monitor present plus last round's score. Ac...
Benchmark Heaven's JevBench scores 534 fixed decisions on four equally weighted axes: intelligence above chance, calibration, speed and cost. Jev 1.13.0 leads at 74.4 for $0.040, SemIf (Qwen3.5-4B) follows at 73.1 for about $0.022, diffusion-Gemma djev at 73.0. GPT-5.6 Luna on...
Fathom targets the scan that ranks all n keys for a top-k step when million-token agent sessions keep KV caches and indexes in host memory. It stores the 4-bit K cache channel-major as bit planes, so a prefix of t planes is exactly that channel's t-bit quantizer, and each quer...
arXiv 2608.04735 points out that monitorability evals overwhelmingly use *explicit* influence, where the prompt tells the model to hide a side task, and monitors catch 60-94% across seven frontier extended-thinking models. Swap in subtle contextual bias and detection drops 41-...
DeepMind's Delegation Capability Tokens paper (arXiv) is the most important agent security paper since the MCP specification. It formally solves the delegation problem: how do agents safely give other agents scoped permissions? The cryptographic caveat system enables least-pri...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.