Fetching from the wire…
Public story · 2026-09-08 · high
SciDocBench, a new benchmark of 124 expert questions, finds AI systems weakest at grounding and verifying evidence.
Why now: SciDocBench's paper was on arXiv as of September 8, 2026.
A new benchmark scored the best multimodal AI system at just 62.6 out of 100 on real scientific reading tasks, per SciDocBench. Document perception, evidence grounding, verification, and cross-document reasoning came in as the weakest categories. Those four skills are what a research assistant is supposed to do: read a paper, pull the right evidence, and check its own answer.
SciDocBench draws on 124 expert-authored questions, screened for difficulty, across seven capability groups and 19 subtasks in five domains. Each question ran under four conditions, English or Chinese, images shown all at once or interleaved with text, producing 496 test instances total.
The systems fell apart in four specific places: reading documents accurately and grounding claims in the right evidence. They also struggled to verify their own answers and reason across more than one document.
Alongside the benchmark, the paper released SciDocIR, a typed evidence-graph format. It preserves document objects, layout, cross-references, and provenance instead of flattening everything into plain text. The release includes about 15,000 supervised training examples built on that structure.
The benchmark doesn't say which architectural fix closes the gap. It just gives the next claim of "we read the paper" something to be measured against.
Each link below shares sources, entities, or timing with this story.
A 6B Diffusion Transformer trained from scratch paired with a frozen VLM understanding module on the LLaDA2.0-Mini backbone, building a visual generative prior through image-only pre-training and mid-training across a 220M-sample pipeline before touching paired image-text data...
Danish Foundation Models trained it from scratch on 161 datasets. Across 20 benchmarks spanning English, math and code, and Danish, it beats the original HRM-Text 1B, sets a new Danish state of the art, and competes with Qwen 3.5 4B and Gemma 4 E2B. Weights are on Hugging Face...
English and Chinese, with reference-free voice design from a natural-language description, reference-guided cloning and low-latency streaming, currently first among open-weight models on the Artificial Analysis TTS leaderboard. The 6GB/4x-realtime figure comes from testers usi...
arXiv 2607.13125: 33 authors, a unified multimodal understanding-and-generation model under Apache 2.0 with weights, code, and recipes, trained on only 208.62 million unique images at a theoretical training cost around $400,000. Four variants (Base, Turbo, Edit, Edit-Turbo) co...
arXiv 2607.27178 releases an end-to-end open recipe against the closed-training-data reproducibility gap: 665M curated English contrastive pre-training pairs distilled from 1.4B across 34 public sources, plus 1.88M supervised fine-tuning pairs with mined hard negatives. Two 14...
Data that contradicts the vibe. That's rare enough to lead with. Dipongkor, Baral, Lam and Moran analyzed 4,882 pull requests from five coding agents in the AIDev dataset (532 Java, 4,350 Python), accepted to ICSME 2026. The findings, in order of how much they should change yo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.