Germán Kruszewski
7 papers in the PaperMetrix corpus
Papers by this author
-
Cooperative Learning of Disjoint Syntax and Semantics
2019 · arXiv (Cornell University)
Serhii Havrylov, Germán Kruszewski, Armand Joulin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.
-
On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
2022 · arXiv (Cornell University)
The availability of large pre-trained models is changing the landscape of Machine Learning research and practice, moving from a training-from-scratch to a fine-tuning paradigm. While in some applications the goal is to "nudge" the pre-trained …
-
Controlling Conditional Language Models without Catastrophic Forgetting
2021 · arXiv (Cornell University)
Machine learning is shifting towards general-purpose pretrained generative models, trained in a self-supervised manner on large amounts of data, which can then be applied to solve a large number of tasks. However, due to their …
-
What you can cram into a single vector: Probing sentence embeddings for\n linguistic properties
2018 · arXiv (Cornell University)
Although much effort has recently been devoted to training high-quality\nsentence embeddings, we still have a poor understanding of what they are\ncapturing. "Downstream" tasks, often based on sentence classification, are\ncommonly used to evaluate the quality of …
-
The emergence of number and syntax units in
2019
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, Marco Baroni. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and …
-
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
2018
Although much effort has recently been devoted to training high-quality sentence embeddings, we still have a poor understanding of what they are capturing. "Downstream" tasks, often based on sentence classification, are commonly used to evaluate …
-
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
2022 · arXiv (Cornell University)
Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed …