GAMUT Measures the Missing Half of Factuality — Best Model Scores 58.7%
Factuality evaluation has focused almost entirely on precision (are the claims correct?) while ignoring completeness (does the answer contain everything it should?). GAMUT introduces a two-level meta-rubric: a structured rubric captures organization and importance of required content, then compiles mechanically into flat binary checks an LLM judge can grade reliably — handling open-ended sets, ordered processes, and inter-fact relationships that flat boolean lists miss. The benchmark spans 1,813 questions grounded in real wearable imagery across 10 domains with expert-verified evidence-backed rubrics; across 14 frontier and open-weight models the best score was 58.7% (Gemini 3.1 Pro), and a text-only variant was also released.
↳ Follow the thread