Researcher profile
Johanna Vielhaben
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
2024 · arXiv (Cornell University)
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their …