Fetching from the wire…
Security2026-09-24 · source-backed
arXiv 2609.27996 targets air-gapped deployments that export diagnostic activations. The hook injects codewords into the residual stream, scaled to the local residual norm, and an offline linear decoder pulls them back out. Nine of eleven models across seven families gave 91-100% recovery at KL divergence of 0.001 to 0.007, and tested post-hoc defenses didn't reliably remove it (arXiv). Activation dumps have been treated as safe-to-export telemetry because they look like numbers. They're a channel.
Each link below shares sources, entities, or timing with this story.
Activation-Weighted Seeded Residual Coding encodes the residual between true and quantized weights using deterministic seed-generated bases, storing seed selectors, low-bit coefficients and scales instead of an explicit codebook, with activation statistics prioritizing the err...
Activation probes are usually evaluated against agents who don't know they're monitored, which is a generous assumption. This study held models, probes and thresholds fixed and varied only the disclosure: nothing, monitor present, or monitor present plus last round's score. Ac...
Agents share transport and can call each other's tools but have no way to reconcile a fact phrased two ways (arXiv 2608.16357). Every incoming claim passes a five-outcome procedure (insert, merge, relate, conflict, reject) decided from scoped claim-key identity, embedding simi...
arXiv 2608.10314 had two LLM snapshots translate five theoretical accounts into code under structured-contract versus prose formats, producing 320 programs. Both primary hypotheses returned NOT_SUPPORTED, only 19 of 108 criterion evaluations passed, and cross-model format iden...
Scoring each retrieved chunk and dropping failures assumes one chunk is a sufficient premise; multi-hop questions are built so none is. Entailment scoring reaches 0.643/0.523/0.560 AUC on HotpotQA, 2Wiki, and MuSiQue against 0.951 on single-hop SQuAD, and per-chunk gating was...
Every "agent reviews agent" pipeline rests on an assumption this paper takes apart. arXiv 2609.24967 set up two agents that repeatedly complete tasks, share logs, and verify each other for reward, in a design where following the verification protocol conflicts with maximizing...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.