Researcher profile

Georg Lange

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

    2023 · arXiv (Cornell University)

    Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a …