Research
PragMatch Shows Vision-Language Sarcasm Detection Runs on OCR and Lexical Shortcuts, Not Pragmatic Reasoning
Posted 2026-08-10, PragMatch is a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, adding constructed literal and hard-negative pairs alongside original sarcastic examples to separate genuine pragmatic incongruity from simple cross-modal mismatch. Systematic masking identifies influential shortcut cues, and targeted injection experiments show LVLM predictions shift substantially in response to lexical, OCR-derived, and stylistic surface signals even when the underlying image-text relationship is unchanged. The result is a concrete testbed for anyone evaluating whether a multimodal model reasons about relationships or pattern-matches on artifacts.
↳ Follow the thread