Research
SIEVES: Selective Prediction Generalizes Through Visual Evidence Scoring for Multimodal LLMs
As vision-language benchmarks approach saturation, real-world deployment requires knowing when the model is wrong, not just being right on average. SIEVES introduces a selective prediction framework for multimodal LLMs that scores visual evidence quality to decide whether to answer or abstain. The method generalizes across tasks without task-specific training, enabling practitioners to set error-rate thresholds for production deployment of vision-language models — critical for any application where wrong answers are more costly than no answer.
Source
↳ Follow the thread