Sources
OpenAI says GPT-Astra plus Codex wrote Jalapeño kernels that beat human-expert code by 1.5-1.8x on attention and MoE blocks
Buried in the same Jalapeño post is a claim that model-assisted kernel work brought three previously unplanned open-weight models to high performance on the new chip in about two months, and that for selected attention and MoE blocks the model-written implementations ran 1.5-1.8x faster than existing human-expert kernels. That moves compiler and kernel engineering inside the model improvement loop rather than treating it as application-layer coding. Treat the numbers as first-party until the full Hot Chips presentation is published, but the direction is the more useful signal: low-level performance work is now a plausible agent task, not just a place agents generate plausible-looking garbage.
↳ Follow the thread