ملف الباحث
Georg Lange
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching
2023 · arXiv (Cornell University)
Mechanistic interpretability aims to understand model behaviors in terms of specific, interpretable features, often hypothesized to manifest as low-dimensional subspaces of activations. Specifically, recent studies have explored subspace interventions (such as activation patching) as a …