Sources
Φ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimization
arXiv 2609.10226 (2026-09-09) is a benchmark built from optimization problems in frontier research and grounded in real code repositories. Tasks range from completing a single kernel function to long-horizon, end-to-end system optimization of the LLM infrastructure stack. The authors say existing benchmarks cover only isolated kernels or preset optimization targets. It is a single-source preprint with no headline numbers in the abstract, but it measures the 'AI improving AI infrastructure' capability that labs keep citing in their self-improvement claims.
Source
↳ Follow the thread