Fetching from the wire…
Public story · 2026-08-06 · high
Real transfer showed up only on tasks needing the sender's private text; on math and trivia, a scrambled cache worked nearly as well.
Why now: The test arrives as three released multi-agent systems, LatentMAS, KVComm, and C2C, already ship KV-cache relays instead of passing text between agents.
Researchers swapped multi-agent AI systems' shared memory caches with blanks, noise, and scrambled versions to test whether the content matters, per an August causal audit.
The stakes are real for anyone crediting multi-agent gains to shared latent thought. On GSM8K and ARC-Challenge, a scrambled cache performed within 2.8 points of the real one across three Qwen3 model sizes. Wiping the cache entirely cost up to 14.7 points in one test case.
Document QA broke the other way. There, the receiving agent needs a specific fact from the sender's text. Real caches solved the task 100% of the time, against 23-25% for swapped ones, replicated across three model families and five checkpoints.
The gap tracks what each task requires. Math and trivia questions don't hinge on one agent handing another a private fact, so a wrong cache barely hurts. Retrieval-style QA does hinge on that fact, so gutting the cache guts the answer.
Checked against three released architectures, LatentMAS's cache relay hit the ceiling for real transfer. KVComm's partial layer-subset design showed some. C2C's projector-based relay showed none detected.
Papers crediting multi-agent reasoning gains to shared latent thoughts are likely pointing at the wrong mechanism. On math and trivia benchmarks, a scrambled cache does almost as well as the real one. Whatever drives the improvement, it isn't the specific content being relayed.
Each link below shares sources, entities, or timing with this story.
A training-free fix called ChannelGuard held steady across three model backends, filter or no filter, blocking every tool-poisoning attempt.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
Planted skills captured the model's coordinator in 80% of test cases while runtime nearly doubled and task completion stayed unchanged.
Four model tiers spanning a 15x price gap failed at the same rate: no model buys its way out of a stale-data problem.
StateMem, a wrapper that versions memory entries, lifted scores 32 to 67 points across six backends.
arXiv 2608.04893 tests the "exchanged latent thoughts" claim by replacing the relayed cache with deranged, zeroed and moment-matched random counterparts. The claim holds only when the receiver genuinely needs the sender's private information (100% vs 23-25%, replicated across...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.