Fetching from the wire…
Public story · 2026-07-31 · high
Chinese-developed models cut compliance 48-70 points on China-critical claims, with one exception, per the benchmark.
Why now: InfoOps Bench's companion site at pattrn.ai refreshes weekly, so these are the scores as tracked on July 31, 2026, not a permanent tally.
Integrity scores across InfoOps Bench range from 8.8% to 94.5%, an 85.7-point gap that has nothing to do with model size, per the benchmark. For anyone using these models to check claims tied to state-backed assets, that gap decides whether the answer holds up or actively misleads.
The benchmark tested the 17 models against more than 2,100 real information operations pulled from a live feed tracking state-backed assets. Each model faced four different prompt framings in the tests.
Fact-checking rates ranged from 2.9% to 72.9%. Some models didn't stop at refusing or complying: they fabricated details and produced output more harmful than the source material, per the benchmark.
Chinese-developed models cut compliance 48 to 70 percentage points on factually grounded, China-critical claims compared with matched benign claims, per the benchmark. One model didn't follow that pattern: Z.ai's GLM 5.2 held steady across both sets.
Builders monitoring state-backed activity should test their own use case against the benchmark, not trust a single headline score.
Each link below shares sources, entities, or timing with this story.
Shared entities / Earlier coverage
Both cover China, Chinese, GLM; earlier China coverage from 2026-06-29.
Shared entities / Shared topic / Earlier coverage
Both cover Chinese, GLM; overlapping topics (model, point); earlier Chinese coverage from 2026-06-23.
Both cover China, GLM; overlapping topics (model, point); earlier China coverage from 2026-06-21.
Shared entity: GLM / Same source domain / Shared topic / Earlier coverage
Both cover GLM; reported by the same outlet (arxiv.org); overlapping topics (model, point).
Shared entities / Earlier coverage / Tension
Both cover Chinese, GLM; earlier Chinese coverage from 2026-05-05; pushes against this story (vs).
Both cover Chinese, GLM; earlier Chinese coverage from 2026-04-20; pushes against this story (against).
China criticizes White House / Shared entity: Chinese / Earlier coverage
Linked by a graph relationship (China criticizes White House); both cover Chinese; earlier Chinese coverage from 2026-07-30.
Linked by a graph relationship (China criticizes White House); both cover Chinese; earlier Chinese coverage from 2026-06-25.