Fetching from the wire…
Research2026-09-17 · source-backed
arXiv 2609.18094 stores autonomous research as an append-only DAG of Git commits, so any agent can check out a prior commit to verify a claim or build on it, with a diversity index preventing population convergence on one approach. A 12-day run on weight initialization produced 1,703 contributions and moved performance from 3.39 to 1.899 bits per byte. The winning solution traced back through 145 commits across 15 accounts, and 165 independent reproductions ran without failure. The reproduction count is what separates this from the usual multi-agent-research demo.
Each link below shares sources, entities, or timing with this story.
Four stories about things going wrong. Here's one about something working, with actual numbers attached. In an August 7 disclosure covered by TechCrunch, Airbnb said AI now writes 60% of its new code, that concept-to-launch time on key initiatives has dropped by as much as 60%...
Efficiency Hallucination is the tendency to issue non-functional mutations with unsubstantiated performance claims on code that's already optimal, driven by binary benchmarks that reward editing over abstaining (arXiv 2609.14839). Across 180 optimization runs on nine GPT, Clau...
Accuracy drops 30–50% well before you hit the documented context limit. Not at the limit. Before it. Cross-model testing across GPT-4.1, the Claude 4 family, Gemini 2.5, and Qwen3 quantified what everyone shipping long-context features has felt and couldn't measure (Glasp). Th...
diegosouzapw/OmniRoute added 1,343 stars on July 20, a single MIT-licensed gateway across 268+ providers (50+ free) and 500+ models including Claude, GPT, Gemini, Kimi K3, GLM and DeepSeek, wired for Claude Code, Codex, Cursor, Cline and Copilot. Quota-aware automatic fallback...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.