Research
AnchorBench: Frontier Models Above 95% Control Accuracy Still Shift Toward Plausible Anchors
AnchorBench evaluates the anchoring bias across fourteen models (ten open-weight, four frontier API) using multiple injection pathways and an explicit anchor-relevance axis separating irrelevant from plausible anchors. Anchoring turns out to be strongly pathway-dependent, plausible anchors move judgments more than irrelevant ones when delivered through stronger pathways, and influence weakens as the anchor moves further from the evidence-supported answer — most visibly on External and RAG pathways. The finding that matters for RAG builders: high accuracy on the anchor-free control (>95% within 10 points of gold) does not predict robustness to a plausible anchor arriving through retrieval.
↳ Follow the thread