Guide Labs Trains Interpretability In Rather Than Bolting It On: Steerling-8B Matches Peers Trained on 2–16x More Compute
arXiv 2608.07594 (Aug 6) from the Guide Labs team, including Andreas Madsen and Julius Adebayo, rejects the standard practice of explaining opaque models post-hoc and instead makes interpretability a training-pipeline constraint optimized jointly with the language modeling objective. Across three orders of magnitude of compute and both autoregressive and diffusion models, they show interpretability improves alongside capability rather than trading against it. The resulting Steerling-8B — a diffusion LM with a causal attention mask — attributes each generated token to input tokens, human-understandable concepts, and training data, and supports closed-loop concept steering with no retraining, while staying competitive with peers trained on 2–16x more compute. It drew 238 upvotes, second-highest on the Aug 11 HuggingFace Daily Papers page.
↳ Follow the thread