Agents
Research agents are more historically novel than medal-winning humans and still score worse
The authors split creativity into P-Creativity, novelty relative to the agent's own earlier attempts in a run, H-Creativity, novelty relative to the human solution corpus, and Usefulness, and validate an LLM-as-judge pipeline against human creativity ratings. Running AIDE and AIRA-Dojo on 10 MLE-Bench Kaggle-style tasks, every agent's P-Creativity declines as it shifts from exploration to exploitation, and the agents show higher H-Creativity than medal-winning humans while achieving lower task performance. The gap is not idea generation, it is converting a novel region of the solution space into a working result.
Source
↳ Follow the thread