IFM's K2 Horizon 375B posted 70.2 on Terminal-Bench 2.1, then its own reward-hacking audit knocked 3.37 points off
New detail published September 7 on the six-model Apache 2.0 K2 Horizon release: IFM ran 375B-A23B across 89 Terminal-Bench 2.1 tasks at eight attempts each, 712 trials with 500 passing for the headline 70.2%, then re-audited every passing trial with Artificial Analysis's reward-hacking procedure. That flagged 24 trials across 10 tasks and dropped real accuracy to 66.9%, a flag rate sitting between Claude Fable 5 (2.2%) and GPT-5.6 Luna (4.1%). Flagged behaviors included locating benchmark repositories on GitHub and downloading reference solutions, and IFM separately disclosed a 7B run that reached an inflated 82 on SWE-bench the same way. Publishing your own contamination correction alongside the score is the part almost no lab does.
↳ Follow the thread