Research
Via Negativa for AI Alignment: Theoretical Account for Why Negative-Only Training Outperforms Preference RLHF
This position paper provides a unified theoretical explanation for the empirical finding that negative-only training (Constitutional AI, Distributional Dispreference Optimization) matches or beats RLHF: positive preferences encode continuously coupled, context-dependent values that lead models to learn surface correlates like sycophancy, while negative constraints are discrete, finite, and independently verifiable, enabling stable boundary learning. The framework — rooted in Popper's falsification logic — proposes shifting alignment from 'what humans prefer' to 'what humans reject' with testable predictions.
↳ Follow the thread