HCIG: Hierarchical Cross-Modal Incongruity Graph Detects Sarcasm and Cyberbullying
arXiv·medium signal
Bhavana Verma, Priyanka Meel and Dinesh Kumar Vishwakarma (arXiv 2607.16076, cs.CV/cs.AI/cs.CL) build a graph network that explicitly models incongruity between image and text — the mismatch that carries the actual meaning in sarcasm and in a lot of harassment content. Flat multimodal fusion tends to average the two modalities together and lose exactly that signal. Directly applicable to trust-and-safety pipelines where text-only classifiers currently miss image-text combinations.