Marc Dymetman
4 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Character-based NMT with Transformer
2019 · arXiv (Cornell University)
Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. …
-
On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
2022 · arXiv (Cornell University)
The availability of large pre-trained models is changing the landscape of Machine Learning research and practice, moving from a training-from-scratch to a fine-tuning paradigm. While in some applications the goal is to "nudge" the pre-trained …
-
Controlling Conditional Language Models without Catastrophic Forgetting
2021 · arXiv (Cornell University)
Machine learning is shifting towards general-purpose pretrained generative models, trained in a self-supervised manner on large amounts of data, which can then be applied to solve a large number of tasks. However, due to their …