Fetching from the wire…
Agents2026-09-03 · source-backed
This paper formalizes an epistemic Sybil problem: another agent is not another observation, because apparently independent reports can descend from the same evidence root and no report-only aggregator can generally distinguish replication from corroboration. Over 20,000 controlled agent report and extraction calls confirm it. Holding one evidence root fixed while report multiplicity rises from 1 to 32 drops naive posterior coverage from 0.940 to 0.263; raising evidence roots from 1 to 16 at fixed report count closes the gap entirely. Replicate extraction errors from a shared base model correlate at gamma 0.719. Fanning out N agents over one source buys confidence, not information.
Each link below shares sources, entities, or timing with this story.
"Reconcile Once, Write Anytime" splits knowledge maintenance from report composition: a deterministic librarian ingests timestamped sources into evidence cards, a metric ledger and claim graphs, and a multi-agent writer may only read evidence stamped at or before a declared cu...
arXiv 2608.04755 injected Android permission popups into real GUI tasks across four frontier multimodal LLMs with synchronized screenshots and UI trees. Holding the task fixed and changing only the requesting app flipped grants from 26/32 to 0/32, an App-Trust Bias. Holding th...
WebMASLab holds task, tools, and browser fixed and varies only architecture. The Telephone Loop attack exploits cross-agent delegation to create cyclical task loops, averaging 80% success with 0% detection against multi-agent versions of Claude Sonnet 4.5, GPT-5.2, and GPT-5.4...
This is the most consequential architecture decision in enterprise software since cloud versus on-prem, and it happened quietly across three vendor announcements. PYMNTS connected the dots first. SAP blocks. Its API Policy v4/2026, published in late April, prohibits using SAP...
The mechanism is copyable and the disclosure is more interesting than the mechanism. Anthropic published on August 31 that it resumed external cybersecurity evaluations after a pause of several weeks, gated behind a real-time classifier that blocks the tool call before executi...
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.