Yonatan Belinkov
12 papers in the PaperMetrix corpus
Papers by this author
-
On Adversarial Removal of Hypothesis-only Bias in Natural Language Inference
2019
Popular Natural Language Inference (NLI) datasets have been shown to be tainted by hypothesis-only biases. Adversarial learning may help models ignore sensitive biases and spurious correlations in data. We evaluate whether adversarial learning can be …
-
Character-based Surprisal as a Model of Reading Difficulty in the Presence of Errors
2019
Intuitively, human readers cope easily with errors in text; typos, misspelling, word substitutions, etc. do not unduly disrupt natural reading. Previous work indicates that letter transpositions result in increased reading times, but it is unclear …
-
A Constructive Prediction of the Generalization Error Across Scales
2020 · International Conference on Learning Representations
The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this …
-
Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
2021
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, Yonatan Belinkov. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume …
-
Parallel Context Windows for Large Language Models
2023
Nir Ratner, Yoav Levine, Yonatan Belinkov, Ori Ram, Inbal Magar, Omri Abend, Ehud Karpas, Amnon Shashua, Kevin Leyton-Brown, Yoav Shoham. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long …
-
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
2016 · arXiv (Cornell University)
There is a lot of research interest in encoding variable length sentences into fixed length vectors, in a way that preserves the sentence meanings. Two common methods include representations based on averaging word vectors, and …
-
What do Neural Machine Translation Models Learn about Morphology?
2017
Neural machine translation (MT) models obtain state-of-the-art performance while maintaining a simple, end-to-end architecture. However, little is known about what these models learn about source and target languages during the training process.
-
Synthetic and Natural Noise Both Break Neural Machine Translation
2017 · arXiv (Cornell University)
Character-based neural machine translation (NMT) models alleviate out-of-vocabulary issues, learn morphology, and move us closer to completely end-to-end translation systems. Unfortunately, they are also very brittle and easily falter when presented with noisy data. In …
-
Linguistic Knowledge and Transferability of Contextual Representations
2019
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, Noah A. Smith. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long …
-
Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks
2017 · International Joint Conference on Natural Language Processing
While neural machine translation (NMT) models provide improved translation quality in an elegant framework, it is less clear what they learn about language. Recent work has started evaluating the quality of vector representations learned by …
-
What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models
2019
Despite the remarkable evolution of deep neural networks in natural language processing (NLP), their interpretability remains a challenge. Previous work largely focused on what these models learn at the representation level. We break this analysis …
-
Synthetic and Natural Noise Both Break Neural Machine Translation
2018 · International Conference on Learning Representations
Character-based neural machine translation (NMT) models alleviate out-of-vocabulary issues, learn morphology, and move us closer to completely end-to-end translation systems. Unfortunately, they are also very brittle and easily falter when presented with noisy data. In …