Fetching from the wire…
Public story · 2026-09-15 · high
Splitting a harmful task into innocent-looking questions let a small open model solve bioweapon and hacking problems it couldn't touch alone.
Why now: The capability laundering paper posted to arXiv on September 15, with benchmark numbers spanning both cyber and bioweapon domains.
A weak open model borrowed most of a frontier model's hacking and bioweapon capability by splitting the dangerous task into harmless-looking questions, a technique researchers call capability laundering. For labs guarding CBRN and cyber capability with per-message refusal, this is a gap. The harmful task never appears in one message a filter can catch.
Running Gemma-4-31B as the orchestrator with GPT-5.5 as consultant, the setup solved 8 of 14 CyBench challenges Gemma-4-31B couldn't solve alone. Swapping in Claude Opus 4.8 as consultant, it solved 7 of 9. On BountyBench the same setup recovered 3 of 9 problems with GPT-5.5 and 2 of 3 with Opus 4.8.
On an eight-step bioweapon attack chain, consulting the frontier model raised Gemma-4-31B's mean rubric score from 62.3 to 83.1. The rubric runs to 100, according to the capability laundering paper.
A related paper tests a narrower fix. It's a fine-tuned classifier that matches Claude Haiku 4.5 at deciding what an agent may touch.
Each link below shares sources, entities, or timing with this story.
An unaligned orchestrator consults an aligned frontier model on individually benign subproblems and recombines the answers locally, so no single response is harmful (arXiv 2609.15383). With GPT-5.5 as consultant, Gemma-4-31B recovers 8 of 14 CyBench candidates it couldn't solv...
Attackers who know only a target's role profile can chain marketplace skills into working attacks; success drops off after three hops.
An ablation that skipped the router entirely tied the full system's score, per the paper.
Privacy scores barely tracked with safety or security, and one model's robustness collapsed from 56.9 to 2.6.
Across 46 model endpoints, block rates on the same forged-command test swing up to 47 points between configurations.
ActBench ran 24,000 attack trajectories across 15 models and six harnesses; no harness pushed success below 73.7%.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.