Research
Diagram-MMU: 12 MLLMs Can Reason About Scientific Diagrams but Can't Turn Them Into TikZ Code
Motivated by OpenAI Prism's diagram-to-TikZ feature, Diagram-MMU offers 3.7k curated diagrams and 18.3k human-validated questions across six domains, testing diagram-to-code parsing, diagram-to-code editing, and diagram QA in both direct and agentic settings. Evaluating 12 MLLMs shows a clean split: models reason well over diagrams but struggle to parse and edit them into code. Agentic scaffolding improves parsing and editing for most models while degrading their QA — Claude-4.6 Opus was the only model that improved consistently across all three tasks.
↳ Follow the thread