Razvan Pascanu
6 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Wide Neural Networks Forget Less Catastrophically
2021 · arXiv (Cornell University)
A primary focus area in continual learning research is alleviating the "catastrophic forgetting" problem in neural networks by designing new algorithms that are more robust to the distribution shifts. While the recent progress in continual …
-
Pushing the limits of self-supervised ResNets: Can we outperform supervised learning without labels on ImageNet?
2022 · arXiv (Cornell University)
Despite recent progress made by self-supervised methods in representation learning with residual networks, they still underperform supervised learning on the ImageNet classification benchmark, limiting their applicability in performance-critical settings. Building on prior theoretical insights from …
-
Architecture Matters in Continual Learning
2022 · arXiv (Cornell University)
A large body of research in continual learning is devoted to overcoming the catastrophic forgetting of neural networks by designing new algorithms that are robust to the distribution shifts. However, the majority of these works …
-
The Tunnel Effect: Building Data Representations in Deep Neural Networks
2023 · arXiv (Cornell University)
Deep neural networks are widely known for their remarkable effectiveness across various tasks, with the consensus that deeper networks implicitly learn more complex data representations. This paper shows that sufficiently deep networks trained for supervised …
-
Round and Round We Go! What makes Rotary Positional Encodings useful?
2024 · arXiv (Cornell University)
Positional Encodings (PEs) are a critical component of Transformer-based Large Language Models (LLMs), providing the attention mechanism with important sequence-position information. One of the most popular types of encoding used today in LLMs are Rotary …
-
Interaction Networks for Learning about Objects, Relations and Physics
2016 · arXiv (Cornell University)
Reasoning about objects, relations, and physics is central to human intelligence, and a key goal of artificial intelligence. Here we introduce the interaction network, a model which can reason about how objects in complex systems …