Fetching from the wire…
Public story · 2026-07-22 · high
Gemini 3.1 Pro leads a 14-model field but only covers 58.7 percent of required content on wearable-image questions, per a new arXiv paper.
Why now: Covered in the 2026-07-22 briefing citing arXiv 2607.19322.
A new benchmark called GAMUT scores AI models on something factuality tests have mostly ignored: whether an answer is complete, not merely whether it's correct. The paper, posted to arXiv, tested 14 frontier and open-weight models against 1,813 questions grounded in real wearable-device imagery across 10 domains, with evidence verified by experts. The best score, from Gemini 3.1 Pro, was 58.7 percent.
That number matters because factuality evaluation has almost entirely measured precision, whether the claims a model makes are true, while leaving completeness unmeasured. A model can state only accurate facts and still fail if it omits half of what a correct answer requires. GAMUT's method is a two-level meta-rubric: a structured rubric that captures how required content should be organized and weighted, then compiled into flat binary checks an LLM judge can grade consistently.
The gap between GAMUT's top score and 100 percent is the story. The strongest model in this test gets less than 60 percent of required content into its answers. That makes completeness a wide-open failure mode that precision-only benchmarks have been hiding, not a minor tuning problem.
Worth watching whether other benchmark builders adopt a completeness axis alongside precision, and whether scores move much as models get optimized against it specifically. If GAMUT's approach holds up, expect current leaderboards to look different once completeness gets factored in. Builders shipping anything that summarizes or answers from images, wearable or otherwise, should treat a correct-sounding answer as no guarantee it's a complete one.
Each link below shares sources, entities, or timing with this story.
Gemini competes with Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, LLM; reported by the same outlet (arxiv.org).
Gemini built by Google / Shared entities / Earlier coverage
Linked by a graph relationship (Gemini built by Google); both cover Gemini, LLM; earlier Gemini coverage from 2026-06-12.
Linked by a graph relationship (Gemini built by Google); both cover Gemini, LLM; earlier Gemini coverage from 2026-05-16.
Meta uses Gemini / Shared entities / Earlier coverage
Linked by a graph relationship (Meta uses Gemini); both cover Gemini, LLM; earlier Gemini coverage from 2026-04-02.
Linked by a graph relationship (Meta uses Gemini); both cover Gemini, LLM; earlier Gemini coverage from 2026-02-12.
Gemini competes with Claude / Shared entity: LLM / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with Claude); both cover LLM; overlapping topics (been, correct).
Gemini competes with Claude / Shared entities / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, LLM; earlier Gemini coverage from 2026-03-03.
Gemini built by Google / Shared entity: Gemini / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini built by Google); both cover Gemini; overlapping topics (been, claim).