Fetching from the wire…
Research2026-07-26 · source-backed
Harvard and Google released the first TPU-native benchmark for AI-generated kernel optimization: 50 JAX workloads, 17 production operators from MaxText architectures (Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, AlphaFold2) and 33 translated from KernelBench at sizes tuned for high TPU v6e MXU utilization (arXiv 2607.20466). Conditioning Gemini 3 Flash on curated Pallas documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at 1.28x geomean speedup. Autocomp's beam search reaches 1.36x over XLA. Target-specific context beats a bigger model, which is the whole argument for skills files in one number.
Each link below shares sources, entities, or timing with this story.
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
"Gemini is Cooked but GCP is Cooking" argues Google quietly shelved 3.5 Pro, which industry chatter placed at roughly Opus 4.5 level, shipping Gemini 3.6 Flash as a bridge the authors call worse than Muse Spark 1.2, Grok 4.5, and tier-1 Chinese open-source models. The hard num...
The changelog offers Z.ai's open-weights coding model with a 1M-token context free via Blackbox AI on AI Gateway, default for new eve agents, switchable for existing ones with eve set --model zai/glm-5.2. Excludes Fast mode and the glm-5.2-fast variant. Separately, Gemini 3.7...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.