Researcher profile
Maximilian Dreyer
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
From attribution maps to human-understandable explanations through Concept Relevance Propagation
2023 · Nature Machine Intelligence
Abstract The field of explainable artificial intelligence (XAI) aims to bring transparency to today’s powerful but opaque deep learning models. While local XAI methods explain individual predictions in the form of attribution maps, thereby identifying …
-
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
2024 · arXiv (Cornell University)
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their …