Fetching from the wire…
Research2026-09-17 · source-backed
arXiv 2609.17865 evaluates an earlier decision point than most safety benchmarks: whether a model chooses to acquire safety-relevant evidence before acting. Across GPT-5.5, o3, Claude Opus 4.8 and Claude Sonnet 4.6, inspection policies differ sharply, with Opus inspecting nearly by default and o3 the most skip-heavy. Inspection rises strongly with severity and falls with retrieval cost, but stated probability barely moves it. A cost-obligation decomposition shows avoidance is driven by retrieval friction and explicit threats to the deployment payoff, not by the duties knowing would create. Telling the model something is risky does much less than making the check cheap.
Each link below shares sources, entities, or timing with this story.
GitHub quietly updated its Copilot pricing multiplier table, and the numbers are jarring. Starting June 1, 2026, every Copilot interaction gets priced in "AI Credits" with per-model multipliers: Claude Opus at 27x the base rate, Claude Sonnet at 9x, and base completions at 1x....
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.