Research
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perturbation and Reward Modeling
Reveals a critical weakness in MLLM judges: when visual evidence conflicts with textual cues, they reward plausible-sounding text over visual accuracy. Proposes perceptual perturbation and reward modeling to recalibrate judge behavior. For anyone using LLM-as-a-judge for multimodal evaluation (image captioning, visual QA scoring), this quantifies how much textual bias corrupts scores and offers a concrete mitigation.
Source
↳ Follow the thread