Research
Trains-but-Doesn't-Learn Benchmark Seats Claude Opus 5, GPT-5.6, Gemini 3.7 Flash and DeepSeek V4-Pro as Forward-Deployed Fine-Tuning Engineers
This benchmark has agents drive a 10-stage post-training delivery pipeline on metered L40S/A100/H200 GPUs over 8B-70B open bases, with an oracle scoring each stage from platform-recorded facts plus a human FDE comparison arm. Its central silent failure is the run where loss falls and every signal stays green, yet the delivered model is no better than base. An operator acceptance gate caught every such run before payment. For anyone letting agents run fine-tuning jobs, the lesson is to gate on held-out lift, not training curves.
Source
↳ Follow the thread