Ryan Lopopolo: your agent is excellent at the concerns you are an expert in and silently bad at everything else
'Aligned to whom?' argues that agent builders lean on model priors for every domain they cannot personally evaluate, so the agent inherits blind spots shaped exactly like the builder's expertise: 'because you are an expert in concerns X, Y, and Z, your agent is likely to be phenomenal at these things...but there are innumerable other concerns that you have either ill- or poorly specified.' Lopopolo extends the problem to evaluation itself, noting the flaw 'generalizes to every auto-rater, every judge, every rubric, every eval, and every researcher,' and calls long-term coherence across agentic work product 'a very unsolved problem.' His conclusion is that alignment is irreducible because 'there is no universal definition of a permissible shortcut.'
↳ Follow the thread