Fetching from the wire…
Agents2026-07-26 · source-backed
Sidik, Levi, and Kimhi argue fixed boot-time topology has no answer when one agent becomes the bottleneck, and propose a six-signal Bottleneck Index (queue depth, context thrash, tool-error rate, role entropy, retry-loop rate, cross-agent wait) that triggers factorizing an overloaded agent into sub-agents while hot-swapping the parent into a coordinator that keeps its external identity (arXiv 2607.20488). Across 720 DeepSeek-V3 runs, the factoriser split raises code-task success from 3.3% to 61.7% and privacy-aware state routing cuts detected high-privacy memory exposure from 2.0 to 0.0 events per task, at under 500µs p99 on the hot path.
Each link below shares sources, entities, or timing with this story.
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
JetBrains expanded Koog with a fluent Java builder API alongside the Kotlin DSL, targeting the enterprise Java ecosystem that Python-first agent frameworks have ignored. Includes Spring Boot integration, multi-provider support (OpenAI/Anthropic/Google/DeepSeek/Ollama), fault-t...
1. Flip your multi-model pipeline to review-then-generate. Instead of using a reasoning model to plan before code generation, let the specialist generate freely and use reasoning tokens for review. Paper shows 90.2% pass@1 vs 87.2% for the planning pattern. Source 2. Audit you...
Show HN: the developer behind JUCE and Cmajor launched an open-source agent where sessions are Yjs-backed CRDT documents instead of chat logs, and nearly everything (context items, loop strategies, slash commands) is a forkable JavaScript plugin. Go plus Wails backend to dodge...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks. Against recognized safety researchers versus general users, Claude dropped behavioral confidence 1.4pp, increased reasoning usage 4.0pp and graded 0.11 points harder. Being tol...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.