Mechanistic Interpretability · EP 1 · Jul 6, 2026
Transformers Are Hiding Their Backup Plans
This week in Mechanistic Interpretability
Papers covered
- Self-repair breaks ablation scoring
- Cheap multi-dim refusal subspaces
- Sparse SAEs via expander graphs
- Diffusion LMs hide a denoising clock
- Steering vectors have hard limits
Your field, every week
Get this for your own field.
Sciport turns the newest papers in your corner of the literature into a short weekly video like this one — auto-curated, delivered to your inbox. First episode’s on us.
Generate your first video →