Data-Science Code Efficiency Does Not Track Correctness: Kimi-K2.5 Ranks Last on Pass, First on Speed
DSEffi-Bench is the first benchmark targeting execution efficiency in LLM-generated data science code, with 1,000 instances across 10-plus libraries, stress-testing harnesses and human-validated references, evaluated on 16 models. GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite solving 47 more tasks, while Kimi-K2.5 is lowest in correctness among frontier models at 40.2% yet posts the best efficiency score at 73.6%. A human-annotated taxonomy finds 79.1% of efficiency deficits come from domain-specific causes rather than algorithmic complexity, and library-conditioned routing approaches Claude-Opus-4.6 Best@3 efficiency at 13.0x lower cost.
↳ Follow the thread