Agents
AgentX-Model: 560 of 636 agent-run production recommender experiments beat their business AUC baseline
arXiv 2609.30001 splits industrial model research into a Research Agent, which writes proposals that are reviewed independently, and a Model Agent, which runs multi-round experiments in sandboxes scoped by business inputs and returns code, measurements and open questions. Work runs through four actions (Reproduce, Follow-up, Composition, Diagnose), and Diagnose handles online feedback such as PCOC prediction bias. In production, 560 of 636 completed model-changing experiments recorded AUC above baseline. It's a concrete loop design for long-running research agents where each run's output picks the next question.
Source
↳ Follow the thread