Mechanist: An Agentic System for Autonomous Interpretability Research, Built on a 13,000-Paper Knowledge Graph Over 43M Papers
Mechanist treats AI as a scientific instrument for discovering the mechanisms behind AI capability, combining an interpretability-focused knowledge graph of ~13,000 papers, a multidisciplinary database of 43 million papers across 26 fields, and a curated library of 32 methods for mechanism analysis, causal intervention, and validation. The 19-author team reports it generates more valuable mechanism hypotheses and executes experiments more reliably than Claude Code and existing AI-scientist systems. Concretely, it uncovered that unsafe traits transfer across modalities through apparently safe training data, developed a mechanism theory of belief formation, and translated those insights into interventions that steer scientific foundation models toward DNA sequences with specified properties.
↳ Follow the thread