Research
IndicSafeEval Finds Jailbreak Success Depends Jointly on Language and Persuasion Style Across Four Indian Languages
LLM safety is still evaluated mostly in English, leaving alignment failures in low-resource and culturally diverse languages unmeasured. IndicSafeEval crosses ten safety-critical content categories with six human-like persuasive strategies across Hindi, Bengali, Marathi and Punjabi, producing 7,200 adversarial prompts for black-box evaluation of several open-source models. Safety behaviour is not uniform: it depends strongly on both the language and the persuasive phrasing, and some risk categories are markedly more susceptible to persuasion-based jailbreaks than others, which means an English-only safety evaluation does not transfer to a multilingual deployment.
↳ Follow the thread