SciDocBench: The Strongest Multimodal System Scores 62.6 of 100 on Real Scientific Reading Workflows
arXiv 2609.05141 introduces SciDocBench, 124 expert-authored and difficulty-screened questions grouped into seven research-assistant capability groups and 19 subtasks across five scientific domains, each instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document representations, giving 496 evaluation instances. The strongest evaluated system reached only 62.6/100, with the weakest areas being document perception, evidence grounding, verification and cross-document reasoning. The authors also ship SciDocIR, a typed evidence-graph representation preserving document objects, layout, cross-references and provenance, plus a roughly 15K-example supervised dataset built on it.
↳ Follow the thread