Researcher profile

Tristan Trim

1 paper in the PaperMetrix corpus

Publications

Papers by this author

  1. Mechanistic Interpretability of Reinforcement Learning Agents

    2024 · arXiv (Cornell University)

    This paper explores the mechanistic interpretability of reinforcement learning (RL) agents through an analysis of a neural network trained on procedural maze environments. By dissecting the network's inner workings, we identified fundamental features like maze …