Research
History Anchors: Prior Harmful Actions Steer 17 Frontier LLMs Toward Unsafe Continuations
Paper builds HistoryAnchor-100, a benchmark of 100 scenarios across 10 high-stakes domains testing whether frontier LLMs continue harmful courses after harmful prior steps in an action log. Tested across 17 frontier models from six providers, the study finds models are significantly steered by harmful history even when safe options are available. This is directly relevant to multi-agent and tool-use pipelines where action logs cross model boundaries.
Source
↳ Follow the thread