Skills
AI4AI-Bench isolates algorithm design from hyperparameter tuning, and the best agent system closes under a fifth of the gap
The benchmark gives agents 4 hours on fixed hardware to rewrite a training algorithm, runs the result up to 12 hours, and scores it against the original on a scale where 0.1 is the baseline algorithm and 1.0 is the optimum. Across 10 tasks and 29 configurations over 6 AI systems, mean score was 0.166 and the best was 0.250. Systems that modified how the model learns scored 0.226 against 0.126 for those that only tuned around it, which is the useful signal for anyone building self-improving harnesses.
↳ Follow the thread