Fetching from the wire…
Models2026-07-30 · source-backed
arXiv 2607.27178 releases an end-to-end open recipe against the closed-training-data reproducibility gap: 665M curated English contrastive pre-training pairs distilled from 1.4B across 34 public sources, plus 1.88M supervised fine-tuning pairs with mined hard negatives. Two 149M models, DenseOn (single-vector) and LateOn (ColBERT-style late interaction), set new size-class state of the art with models, datasets, and training code released. The multilingual finding is the useful one: after translating to eight languages into 307M mmBERT-based variants, the dense model degrades outside translate-train support while late interaction generalizes to unseen languages and scripts.
Each link below shares sources, entities, or timing with this story.
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
Hugging Face's August 26 post introduces ColBERT-style late-interaction training with MultiVectorEncoder, MultiVectorEncoderTrainer, CachedMultiVectorMultipleNegativesRankingLoss and a MultiVectorInformationRetrievalEvaluator (Hugging Face). Their finetuned mLateOn-medical rea...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Thirty-five technique papers tested against the simplest alternative: one auto-generated prompt on a newer-generation model, no iterative refinement (arXiv 2609.00468). Constructive techniques like code generation and repair are the most substitutable. A surviving set relies o...
LangChoiceBench covers 28 projects across seven software areas where Python is a poor default, run against 25 LLMs. Python stays heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models show stronger bias. Analysis of 9,826 reason...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.