Fetching from the wire…
Public story · 2026-09-20 · high
The paper models judgment as a neuron and finds wrong consensus becomes impossible below a 6.4-message read average.
Why now: The paper posted to arXiv this September, at a point when tunable thresholds like this one are still rare in multi-agent debate research.
Multi-agent AI debate reduces to one design knob, how many of the other agents' messages each one reads, according to a paper posted to arXiv.
The finding matters for anyone chaining agents together to vote on an answer, since debate is supposed to filter out mistakes, not spread them. Reading too many other agents' messages can let a wrong answer take over the whole swarm instead.
The setup treats each agent's judgment as a stochastic binary neuron. It's a logistic function fed by a weighted, divisively normalized sum of the messages sitting in its inbox.
Across 31,824 randomized queries with an 8B model, the researchers computed wrong-consensus reachability directly from the network's weights and degree statistics. Below an average of 6.4 messages read per agent, it wasn't reachable from any starting state.
Most multi-agent debate designs pick a communication pattern by feel and hope it converges on the right answer. This paper gives builders something to check instead: compute the message-read average for a topology and see if it clears 6.4. The test so far covers one 8B model on synthetic queries, not the larger models or messier real-world questions most debate systems actually run on.
Each link below shares sources, entities, or timing with this story.
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
CiteShade turns the audit mechanism into the attack surface (arXiv 2609.15660). An attacker controlling a single source gets the system to produce an attacker-chosen wrong answer, attribute it to a trusted source that doesn't support it, and leave the correct evidence retrieve...
ERPBench evaluates six screenshot-only computer-use agents against a live reproducible ERP system, scoring against ground-truth database values rather than screen state. Strong general GUI performance does not transfer. The agents reach the right form and save it; the stored r...
Simulating HBM, DRAM and SSD tiers against a random-forest execution-time predictor across chat, agent and document QA workloads, tiering supported 73.02x more concurrent sessions per GPU at 62.04x lower cost per session. The authors attribute the gains to tier capacities of 1...
arXiv 2609.13356, 227 HuggingFace upvotes, releases a 7B dense model trained from scratch on the premise that small models can't memorize the web but can trade parametric capacity for deliberate thinking plus external tools. Fully open recipe: interleaved gated sliding-window...
Thirteen authors ran a generational genetic algorithm over specialized agents that separately handle mechanistic argument, assumption reconsideration, and evidence and testability assessment (arXiv 2609.15938). Evaluated against DepMap and Open Targets across 34 cancer types,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.