Fetching from the wire…
OSS2026-09-23 · source-backed
Benchmark Heaven's JevBench scores 534 fixed decisions on four equally weighted axes: intelligence above chance, calibration, speed and cost. Jev 1.13.0 leads at 74.4 for $0.040, SemIf (Qwen3.5-4B) follows at 73.1 for about $0.022, diffusion-Gemma djev at 73.0. GPT-5.6 Luna on low effort posts the top intelligence score at 95 but ranks 14th at $0.242. Harness and public tasks MIT-licensed, results JSON published with a sha256, so you can rerun it on your own routing before paying for anything.
Each link below shares sources, entities, or timing with this story.
GitHub published Project HydraFusion on September 4. Spotify published Portal on September 3. CodeRabbit published its Astra evaluation on September 4. None of them coordinated, and all three are the same argument. HydraFusion is a Copilot research preview that treats workflow...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Vicki Boykis wrote a post titled exactly that, "Running local models is good now," and it hit 1,437 points on Hacker News with 551 comments. Her claim is specific and checkable. Gemma 4, the gemma-4-26b-a4b and gemma-4-12b-qat variants, runs agentic coding at roughly 75% of fr...
xAI released Grok 4.7 on September 21 on a new 2.1-trillion-parameter base, a 40% jump over Grok 4.6's 1.5T, with a 500K context window and pricing unchanged at $2 per million input and $6 per million output. Published scores: 46.3% on CursorBench 4.0 (up from 40.4%), 71.0% on...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.