Fetching from the wire…
Research2026-09-22 · source-backed
The paper ties the failure to one quantity: top-1 survives quantization only when the gap between the two highest scores exceeds twice the largest rounding error. Classification loss functions push the correct logit away from the rest, so most inputs clear that bar. Retrieval has nothing separating the top document from the second, so few queries do. The gap needs no labels, which makes it usable both pre-deployment to predict which models will break and per-input at serving time to predict whether the quantized answer matches full precision. (arXiv 2609.24322)
Each link below shares sources, entities, or timing with this story.
AGENTQ is the first study of this attack against agents rather than free-text generation, where the payload is a structured function call nobody reads (arXiv 2609.14060). Naive adaptation of prior backdoor methods wrecks benign utility; AGENTQ combines layer-banded LoRA inject...
The argument is that a structurally compressed model's bfloat16 checkpoint is itself only a distillation-recovered approximation, so training the 4-bit student against it inherits that error. Distilling directly from the original model reaches a comparable peak about 7x faster...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
arXiv 2609.18052 had Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5 solve 992 algorithmic problems as Java Spring Boot service methods against a mandated signature and DTO spec, iteration forbidden, hardcoded answers banned, producing 7,593 methods and 7,936 measured reques...
arXiv 2609.20045 audits context compression with paired histories that share the same current answer, receive the same future update, then require different answers. A deterministic frontier selector scored 96/96 strict reveal accuracy on DeepSeek but 82/96 on GLM, a structure...
Across four benchmarks and 36 judge-examinee pairs, a model's own task accuracy predicts its judging accuracy at r ≥ 0.90, but capability doesn't buy fairness: more capable examinees get more lenient judgments from every judge at r ≥ 0.83. Calibrated weighted majority voting e...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.