Load Hijack Poisons Only MoE Router Weights — 92.3–95.6% of Triggered Tokens Land on One GPU, Cutting Throughput to 0.86x
A malicious model provider modifies nothing but a checkpoint's router weights, keeps a private trigger, and when that trigger appears the router concentrates token-to-expert assignments onto experts co-located on a single GPU, making it a straggler while peers idle. A three-stage optimization resolves the conflict between trigger-dependent concentration and near-clean routing on ordinary inputs; across three MoE families and four corpora the attack directs 92.3% to 95.6% of triggered assignments to target experts. In live expert-parallel serving, triggered traffic produces 1.43x time-to-first-token and 0.86x throughput — a supply-chain attack on the serving schedule itself, invisible to output-quality checks.
↳ Follow the thread