ملف الباحث

Sebastian Lapuschkin

4 أوراق في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Shortcomings of Top-Down Randomization-Based Sanity Checks for Evaluations of Deep Neural Network Explanations

    2022 · arXiv (Cornell University)

    While the evaluation of explanations is an important step towards trustworthy models, it needs to be done carefully, and the employed metrics need to be well-understood. Specifically model randomization testing is often overestimated and regarded …

  2. From attribution maps to human-understandable explanations through Concept Relevance Propagation

    2023 · Nature Machine Intelligence

    Abstract The field of explainable artificial intelligence (XAI) aims to bring transparency to today’s powerful but opaque deep learning models. While local XAI methods explain individual predictions in the form of attribution maps, thereby identifying …

  3. PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits

    2024 · arXiv (Cornell University)

    The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically and encode for multiple (unrelated) features, which renders their …

  4. Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

    2025 · arXiv (Cornell University)

    Robust tooling and publicly available pre-trained models have helped drive recent advances in mechanistic interpretability for language models. However, similar progress in vision mechanistic interpretability has been hindered by the lack of accessible frameworks and …