Fetching from the wire…
Public story · 2026-09-06 · high
The model predicts a tipping point past which habitual chatbot use becomes hard to reverse, untested against real data.
Why now: The paper's arXiv ID places it in September 2026, and it's still drawing debate on Hacker News as of September 6.
A new paper models habitual LLM use like an infectious disease, complete with transmission rates and a tipping point. Past a critical threshold, LLMs as a Cognitive Virus argues that dependence stops being a choice and becomes something the system reinforces on its own. For anyone designing products around daily chatbot habits, that implies a line where encouraging more use stops being reversible.
The framework sorts users into three states: uncoupled, coupled, and persistently dependent. The paper says social transmission, seeing others rely on a chatbot, combines with collective reinforcement to push adoption past that threshold. After that point, the dynamics run on their own regardless of individual choice, the paper argues.
The same model produces what the authors call cognitive immunization: cut transmission and preserve reversibility, and the runaway state doesn't happen. None of this comes from measured behavior. The paper has no empirical component, so the threshold is a hypothesis, not something anyone has observed in real chatbot use.
Each link below shares sources, entities, or timing with this story.
LLMs barely correct errors in their own reasoning traces but readily correct the identical claim when it's attributed to an external source. Relabeling a claim from the agent's own role to an external one raises the explicit-correction rate by 23 to 93 percentage points across...
First benchmark of off-the-shelf LLMs against expert-derived ground truth built on INCOSE criteria, ten models across two families and five generations each, one hundred independent runs, two requirement sets, five temperatures. The error profile is asymmetric, and performance...
This comparison ran the baseline the retrofitted-linear-attention literature skipped. Across multiple LLMs and downstream tasks SWA with sinks matches or beats post-trained linear attention, and on Needle-in-a-Haystack and BABILong it scores 2 to 10 times higher. The recommend...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.