Fetching from the wire…
Agents2026-07-28 · source-backed
arXiv 2607.22465 argues per-call routing is structurally mismatched to agentic work, since long-horizon workflows only produce a delayed task-level outcome and per-call routers can't attribute feedback to individual decisions. It assigns each task to a backend once at admission via contextual bandit, pins every subsequent call, and updates on terminal reward weighting accuracy and latency jointly. tau2-Bench: 7-8 points over latency-matched interpolation. Terminal-Bench: 7.1 points above the strongest single-model baseline with 36% lower latency. (arXiv 2607.22465)
Each link below shares sources, entities, or timing with this story.
Tencent-affiliated work formalizing hosted model routing as a stochastic process (arXiv 2608.16391). Reports average fidelity loss from repeated frozen-context requests and extreme fidelity loss from the upper tail of long-sequence runs, with AFL showing strong linear agreemen...
Quesma ran the model across GPQA Diamond, IFBench and Terminal-Bench 2.1 (89 agentic coding tasks) on L40S, H100 and H200 via Modal. Q4_K_M at 17 GB matched BF16 at 55 GB within a point on all three. UD-Q2_K_XL at 10.7 GB held instruction-following but dropped Terminal-Bench f...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
Here's the number that should sit in every "agents will replace engineers" thread: 15.2%. That's the best model. The mean across 15 frontier models is 4.3%. Tencent Hunyuan's Long-Horizon-Terminal-Bench put 15 frontier models against 46 long-horizon terminal tasks across nine...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.