A Deployed On-Device Model Confabulates on 69% of False Premises While Refusing 18% of Benign Prompts, With No Detectable Signal
This reliability audit targets the developer-accessible on-device foundation model shipping to hundreds of millions of devices with no server-side moderation, asking whether a resource-constrained developer can tell when it is wrong. It finds task-asymmetric miscalibration, guardrails failing in opposite directions across tasks, atop self-reported confidence that is saturated and non-discriminative (AUROC 0.47, ECE 70, worst among comparable small models). Critically, confident-correct and confident-wrong outputs are surface-indistinguishable: a classifier over 15 user-visible features separates them at AUROC 0.55, and no cheap single-generation signal exceeds 0.68. A black-box consistency wrapper requiring no model access cuts confident confabulation from 75% to 3% and lifts selective accuracy from 43% to 83%.
↳ Follow the thread