The Last Translation Benchmark Ships Handcrafted Verification Rules Instead of a Score
Standard machine translation benchmarks are approaching saturation while automatic metrics are unreliable and vulnerable to reward hacking, and gold human evaluation often lacks reproducibility, objectivity and scalability. The Last Translation Benchmark is a collection of human-authored, peer-reviewed examples spanning text, images, audio and video that break leading MT models, and its evaluation approach is the notable part: each example carries handcrafted verification rules describing the concrete failure case on that example, so assessment is actionable rather than a single opaque score. It is a live dataset accepting ongoing contributions, with LTBv1 containing accepted submissions prior to 1 September 2026 and further releases planned.
↳ Follow the thread