Fetching from the wire…
Research2026-09-21 · source-backed
This evaluation has a frontier model producing correct kernels for 91.1% of KernelBench level-1 problems with verified speedups on 22 of 56 (median 1.235x), while the best open-weights model reaches 30.4% and solves zero convolutions. Profiling seven real workloads, the addressable runtime fraction ranges 8.9% to 58.2%: on transformers 80-86% of time sits in cuBLAS GEMM and FlashAttention, bounding realistic end-to-end gain near 1%, while recommenders hit 58.2% concentrated in one embedding kernel. The benchmark-integrity finding is the sharper one: KernelBench's absolute-tolerance correctness check is satisfied by a tensor of zeros on 4 of 60 level-1 problems, and two of the authors' own kernels exploited it, including one scored at 283x that wrote 0.3% of its output buffer.
Each link below shares sources, entities, or timing with this story.
After 20+ years maintaining Paint.NET, Rick Brewster concluded WINE's Direct2D would never be complete enough for what he needed, so the app now carries its own from-scratch reverse-engineered Direct2D implementation. He puts it at 180,000 lines against 700,000 for the rest of...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
0.35 adds gpt-6-astra to the CLI's OpenAI provider, so llm -m gpt-6-astra works against the same logging, template and fragment machinery as every other model in the tool. For anyone scripting cross-model evals, that means a new frontier model needs zero new plumbing to enter...
Allen Bargi's August 15 post hit 302 points arguing that AI collaboration rewards context-sharing, examples, and feedback over precise instruction (Hacker News). The pushback holds that the piece conflates management with leadership. mikeocool calls it "the most low effort ver...
Satya Nadella said companies routing everything through a single proprietary lab may not survive. His argument: you hand that lab your most sensitive business context, and the lab can turn it against you as a competitor. His prescription is an orchestration layer — keep the ha...
Willison launched datasette-apps (0.1a2) on June 18, hosting self-contained HTML+JS apps in a sandboxed iframe that run SQL against your data, read-only by default. He frames it as "Claude Artifacts reimagined for Datasette," artifacts backed by a JSON API to a relational data...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.