Neel Nanda
3 papers in the PaperMetrix corpus
Papers by this author
-
Progress measures for grokking via mechanistic interpretability
2023 · arXiv (Cornell University)
Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that …
-
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
2023 · arXiv (Cornell University)
Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a …
-
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
2025
Model diffing is the study of how fine-tuning changes a model’s representations and internal algorithms. Many behaviors of interest are introduced during fine-tuning, and model diffing offers a promising lens to interpret such behaviors. Crosscoders …