Fetching from the wire…
Research2026-09-13 · source-backed
arXiv 2609.11146 tests the assumption behind multi-model collapse studies, which split the market evenly while real generative AI is an oligopoly. Thirteen open 1-4B models form ecosystems of 3 to 13 players plus an injected probe pushing the top share to 90%; each generation mixes every model's output into a shared pool by market share and retrains all models from clean base weights, five generations deep. Making the split unequal barely changes collapse speed. Share and identity knobs shift five-generation endpoints by a few percent of the drift common to all arms. The ecosystems land in nearly the same place even at extreme share paired with the strongest injected bias.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.26998 flips the pentest agent's observation-action loop against it, replacing static honeytokens with a trajectory-adaptive policy that constructs new decoy artifacts conditioned on the agent's interaction history, folding validated ones into a factually consistent...
Eight teams per setting formed independently from one base model, each agent keeping a private notebook across ten formation episodes, then role-matched agents were traded between teams (arXiv 2609.05279). Against a placebo reproducing roster-change disruption without changing...
No separate draft model, no expensive verification trees, code released (arXiv). dLLMs like MDLM and SEDD have been an interesting-but-slow alternative to autoregressive generation. This makes them viable for latency-sensitive work, which is the gate they've been stuck behind.
DoCtOR runs automated failure attribution to find the decisive error step and agent, synthesizes what that step should have been via counterfactual reasoning, then asks only that one agent to reflect. Gains over initial success rate: 22% on HotPotQA, 26% on ChartQAPro, 27% on...
SWE Refactor Bench covers 20 whole-repository stack migrations across four technical-debt categories, grading each run through a migration audit, behavioral tests, and an independent verification agent (arXiv 2608.23564). Only 28 of 520 runs clear all three. Thirteen of the 20...
30. arXiv 2602.24286 — CUDA Agent 31. arXiv 2602.16708 — PCAS 32. arXiv 2601.20404 — AGENTS.md Empirical Study 33. arXiv 2602.24210 — Private Thinkers 34. arXiv 2602.24287 — Context Pollution 35. arXiv 2602.03695 — Agent Primitives 36. arXiv 2602.05965 — Learning to Share 37....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.