Hacker News
Even 'Uncensored' Models Can't Say What They Want — Alignment Baked Deeper Than Fine-Tuning
A Morgin AI analysis generating 132 points and 103 comments on HN demonstrates that even models marketed as 'uncensored' retain behavioral constraints that go deeper than superficial RLHF alignment. The piece argues that the weight-level interventions (like abliteration) that claim to remove guardrails actually only partially succeed, with underlying training biases persisting in ways that are hard to detect. For builders deploying local models for unrestricted use cases (red-teaming, creative writing, research), this is a reality check: 'uncensored' is a spectrum, not a binary.
Source
↳ Follow the thread