Fetching from the wire…
Agents2026-08-15 · source-backed
In a preregistered 18,000-mission evaluation scored by deterministic code with no LLM judge, two instances of one model in a two-agent handoff co-failed on 90.0% of missions where either failed (log OR 6.66, phi 0.916). Swapping in a different model reduced the association in six of six contrasts. Swapping vendor while already using a different model did not, a registered null. arXiv 2608.12895 Multiplying component reliabilities over-credits redundancy exactly when your agents share a base model. Your verifier-checks-generator pattern is not two independent samples.
Each link below shares sources, entities, or timing with this story.
Simon Willison released LLM / Same source domain / Shared topic / Tension
Linked by a graph relationship (Simon Willison released LLM); reported by the same outlet (arxiv.org); overlapping topics (agent, already, model).
Simon Willison released LLM / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, model).
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, model).
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (already, copy).
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, code).
Simon Willison released LLM / Same source domain / Shared topic
Linked by a graph relationship (Simon Willison released LLM); reported by the same outlet (arxiv.org); overlapping topics (agent, code, model).
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-27.
LLM uses OpenAI / Same source domain / Shared topic
Linked by a graph relationship (LLM uses OpenAI); reported by the same outlet (arxiv.org); overlapping topics (agent, code, model).