ClinHallu: a benchmark for stage-wise hallucination in medical MLLM reasoning
arXiv·medium signal
ClinHallu (arXiv 2606.14697, cs.CV/cs.AI/cs.CL) introduces a benchmark that diagnoses where in a multi-step reasoning chain a medical multimodal LLM hallucinates, rather than just scoring final answers. The stage-wise framing is the useful generalizable idea — locating the reasoning step where a model goes wrong is applicable beyond clinical use to any high-stakes MLLM pipeline.