Research
KoNA Finds VLMs Fail Hardest When a Query Mixes Answerable and Unanswerable Parts
arXiv 2609.04720 argues existing non-compliance benchmarks wrongly assume each request is entirely answerable or entirely refusable, when real queries mix both. KoNA evaluates selective non-compliance across five categories, False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility and Safety, testing both query-level and component-level non-compliance under paired single and compound queries. Across diverse VLMs, models frequently failed to refuse, correct or abstain appropriately, and failures got worse specifically in the compound case where partial compliance is the right answer.
↳ Follow the thread