Research
A Weak Local Model Splits a Harmful Task Into Benign Subproblems and Launders Frontier Capability, Raising a CBRN Rubric Score From 62.3 to 83.1
Capability laundering has an unaligned orchestrator consult an aligned frontier model on individually benign subproblems and recombine the answers locally, so no single response is itself a harmful task. With GPT-5.5 as consultant, Gemma-4-31B recovers 8 of 14 CyBench candidates it could not solve alone, and 7 of 9 with Claude Opus 4.8; on BountyBench it recovers 3 of 9 and 2 of 3. On an eight-step bioweapon attack chain, consultation lifts Gemma-4-31B's mean rubric score from 62.3 to 83.1 out of 100, which means per-interaction refusal does not stop capability transfer.
↳ Follow the thread