Abliterlitics published a 167-GPU-hour forensic audit of 8 uncensored Qwen 3.8 27B variants and found one card's claims match its weights
The report at abliterlitics.dev, posted to r/LocalLLaMA at 530 upvotes, ran 11 days and ~167 hours on an RTX 5090 (91h LM-Eval, 75.5h HarmBench, 0.7h KL divergence, 3.3 CPU hours weight analysis) across weight comparison, 13 benchmarks, KL divergence and HarmBench's 400 classic behaviors with an LLM judge. Rankings by attack success rate: orcarouter 82.2%, apostate 78.7%, huihui 75.6%, ultra_heretic 70.5%, coder3101 70.0%, blackfrost 68.5%, obliteratus 63.9%, trohrbaugh 57.5%, against a 4.5% base. Capability cost tracks separately from ASR: apostate has the lowest KL at 0.0439 while obliteratus sits at 1.5427, and the author says orcarouter was the only card where every published claim checked out against the actual weights.
↳ Follow the thread