ملف الباحث

Neel Nanda

3 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Progress measures for grokking via mechanistic interpretability

    2023 · arXiv (Cornell University)

    Neural networks often exhibit emergent behavior, where qualitatively new capabilities arise from scaling up the amount of parameters, training data, or training steps. One approach to understanding emergence is to find continuous \textit{progress measures} that …

  2. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

    2023 · arXiv (Cornell University)

    Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a …

  3. Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning

    2025

    Model diffing is the study of how fine-tuning changes a model’s representations and internal algorithms. Many behaviors of interest are introduced during fine-tuning, and model diffing offers a promising lens to interpret such behaviors. Crosscoders …