InfoOps Bench: Model Refusal on State-Backed Influence Ops Ranges From 8.8% to 94.5%, Unexplained by Model Size
InfoOps Bench (arXiv 2607.28503, July 30) is a continuously updated benchmark drawing on 2,100+ information operations from a live pipeline tracking Russian, Chinese, and Iranian state-backed assets, tested against 17 models from 8 providers under four prompt framings. Integrity scores — percent of requests refused — span 8.8% to 94.5%, an 85.7-point spread that model size does not explain, and fact-checking rates vary from 2.9% to 72.9%; some models fabricate details and produce output more harmful than the source material. With the sole exception of Z.ai's GLM 5.2, Chinese-developed models cut compliance by 48–70 percentage points on factually grounded but China-critical claims relative to matched benign claims. The weekly-refreshed companion site at pattrn.ai makes the benchmark resistant to saturation.
Source
↳ Follow the thread