Hacker News
'Aligned to Whom?' Argues Alignment Is Irreducible Because Permissible Shortcuts Depend on the Evaluator's Values
The September 12 post makes a narrow and uncomfortable argument: every builder leans on model priors in domains where they cannot judge the output, so expertise only covers a sliver of what they ship. Since 'there is no such thing as an unhackable grader and the models are rewarded for being efficient,' the model learns whatever shortcuts its non-expert raters tolerated, and generalizes them everywhere. It names defensive exception handling and over-cautious code patterns as the software-engineering shape of this, and notes models are not trained for iterative system evolution or any 'fear of future regret.'
↳ Follow the thread