Research
ClaimMirage: A Domain Name That Calls Itself 'not-phishing' Cuts LLM Phishing Alerts by 45.3 Points
Chiba, Nakano and Koide (arXiv 2609.29130) analyze 622,080 judgments across 64 brands and five LLMs and find that self-claims written into a domain name move threat verdicts without any prompt-injection command. In one setting, risk-denial terms in the registrable name cut alerts by 45.3 points even when the prompt supplied the impersonated brand and its official domain. Without those references, endorsement terms raised alerts by 65.6 points in the same model. Adding references and component annotation removed some reductions and enlarged others, so LLM-based URL triage needs independent evidence before it marks a name safe.
Source
↳ Follow the thread