Researcher profile
Georg Lange
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
2023 · arXiv (Cornell University)
Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a …