DenseOn and LateOn: Fully Open 149M Retrievers Set Size-Class SOTA at 56.20 and 57.22 nDCG@10 on BEIR
arXiv 2607.27178 (2026-07-29) releases an end-to-end open recipe against the closed-training-data reproducibility gap: 665M curated English contrastive pre-training pairs distilled from 1.4B across 34 public sources, plus 1.88M supervised fine-tuning pairs with mined hard negatives. The result is two 149M-parameter models — DenseOn (single-vector) and LateOn (ColBERT-style late interaction) — at 56.20 and 57.22 average nDCG@10 on BEIR, new state of the art for that size class, with models, datasets, and training code released. The multilingual finding is the useful one: after translating to eight languages (2.8B pairs) into 307M mmBERT-based variants, the dense model degrades outside translate-train support while late interaction generalizes to unseen languages and scripts.
↳ Follow the thread