Aligning Latent Moral Representations Instead of Responses Improves Adversarial Robustness Where Behavioral Alignment Made It Worse
arXiv 2609.04022 tests whether LLMs preserve prototype-style graded categorization of moral concepts and finds that across 23 models they often fail to distinguish opposed moral categories or preserve fine-grained typicality, with the deficit persisting across parameter sizes and alignment stages. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended judgements at the response level while leaving categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Representational similarity optimization, which supervises latent representations rather than generated responses, produced more modest gains on explicit judgements but consistently improved adversarial robustness across model scales, benchmarks and attack strategies.
↳ Follow the thread