Research
HalluTruthQA-4K Releases 4,000 Arabic Hallucination Instances With Character-Level Error Spans and Human Explanations
Most Arabic hallucination resources label a whole response as hallucinated or not, giving no signal about which content is wrong or why. This expanded corpus covers 4,000 expert-curated QA instances across Islamic knowledge, history, science, and geography — 1,643 hallucinated and 2,357 clean — each pairing a question with a model response, verified reference answer, and five plausible distractors, with 1,843 annotated erroneous spans plus human-written explanations and hierarchical hallucination types. It is the official dataset for Track 2 of the HalluScoring 2026 shared task, and supports span-level error localization rather than binary detection.
↳ Follow the thread