OSS
Heretic automates refusal-ablation for open models and claims 3/100 refusals with 0.16 KL divergence, 32,148 stars and 5,000+ derived models published
Heretic (p-e-w/heretic) combines directional ablation with TPE-based parameter search via Optuna to orthogonalize the residual directions that encode refusal, and it does the parameter tuning automatically rather than by hand. On Gemma-3-12B-IT the README reports 3/100 refusals on harmful prompts versus 97/100 for the original, at a KL divergence of 0.16 on harmless prompts against 0.45-1.04 for competing abliteration approaches. It hit 254 points on Hacker News on 2026-09-21, sits at 32,148 stars and 3,611 forks, supports dense and MoE transformers, and processes a 4B model in 20-30 minutes on an RTX 3090.
Source
↳ Follow the thread