Dispatch
Hawkeye gets coding agents to write GPU kernels 18.9x faster than expert Triton on new attention variants
Import AI 470, published 2026-08-24, covers Hawkeye from Harvard, Stanford, Together AI and Caltech, a framework that lets coding agents produce hardware-aware GPU kernels with minimal expert supervision by steering them with curated unit tests. It reports an 18.9x geomean speedup over expert-written Triton kernels on emerging attention variants, with much smaller gains on settled workloads at 1.22x on Blackwell and 1.00x on MI350, across NVIDIA Ampere, Hopper, Blackwell and AMD MI350. The gradient across those numbers is the actual finding: agent-written kernels win where no human has yet spent months tuning, not where they have.
Source
↳ Follow the thread