Researcher profile

Johanna Vielhaben

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits

    2024 · arXiv (Cornell University)

    The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their …