Fetching from the wire…
Research2026-09-14 · source-backed
Across four benchmarks and 36 judge-examinee pairs, a model's own task accuracy predicts its judging accuracy at r ≥ 0.90, but capability doesn't buy fairness: more capable examinees get more lenient judgments from every judge at r ≥ 0.83. Calibrated weighted majority voting estimates each judge's error rates purely from inter-judge disagreement, needing no ground truth, and stays within 0.5pp of an oracle under distribution shift.
Each link below shares sources, entities, or timing with this story.
First benchmark of off-the-shelf LLMs against expert-derived ground truth built on INCOSE criteria, ten models across two families and five generations each, one hundred independent runs, two requirement sets, five temperatures. The error profile is asymmetric, and performance...
Translating key-value state from one model into a form another can consume works across scale, architecture, attention configuration, tokenizer and family (arXiv 2608.30963). Llama3.1-70B to Qwen2.5-7B reaches 44.0% accuracy against 45.7% native while dropping latency to 138ms...
A synthetic benchmark constructs conflicts where exactly one evidence source matches ground truth, independently varying modality, recency, stated reliability, and provenance. Across open-weight instruction-tuned models the arbitration is systematic: distinct text-versus-numbe...
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
This study investigates dropping speculative decoding's lossless guarantee without any training, quantifying speed-ups against controlled capability drift. Standard spec decoding exactly preserves the sampling distribution. Relaxing it buys latency at a small distributional co...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.