arXiv:2608.07594v1 Announce Type: new Abstract: Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale. We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.
Scaling Inherently Interpretable Language Models
A new arXiv paper challenges the notion that interpretability comes at the cost of capability, showing that making interpretability a training constraint allows it to scale with model capability across three orders of magnitude of compute. The authors introduce Steerling-8B, a diffusion language model with a causal attention mask, which attributes outputs to input tokens, concepts, and training data, enabling closed-loop intervention without retraining. Steerling-8B remains competitive with open peer models trained on 2-16x more compute, suggesting a new scaling paradigm where interpretability is designed into training and improves with scale.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.