Omer Levy
14 papers in the PaperMetrix corpus
Papers by this author
-
Improving Transformer Models by Reordering their Sublayers
2020
Multilayer transformer networks consist of interleaved self-attention and feedforward sublayers. Could ordering the sublayers in a different pattern lead to better performance? We generate randomly ordered transformers and train them with the language modeling objective. …
-
What Do You Get When You Cross Beam Search with Nucleus Sampling?
2021 · arXiv (Cornell University)
We combine beam search with the probabilistic pruning technique of nucleus sampling to create two deterministic nucleus search algorithms for natural language generation. The first algorithm, p-exact search, locally prunes the next-token distribution and performs …
-
Improving Distributional Similarity with Lessons Learned from Word Embeddings
2015 · Transactions of the Association for Computational Linguistics
Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain …
-
Do Supervised Distributional Methods Really Learn Lexical Inference Relations?
2015
Omer Levy, Steffen Remus, Chris Biemann, Ido Dagan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments
2017
While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set (sentence IDs) …
-
Annotation Artifacts in Natural Language Inference Data
2018
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short …
-
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
2018 · arXiv (Cornell University)
For natural language understanding (NLU) technology to be maximally useful, both practically and as a scientific object of study, it must be general: it must be able to process language in a way that is …
-
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
2019 · arXiv (Cornell University)
In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, …
-
SpanBERT: Improving Pre-training by Representing and Predicting Spans
2020 · Transactions of the Association for Computational Linguistics
We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the …
-
Ultra-Fine Entity Typing
2018
We introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g. skyscraper, songwriter, or criminal) that describe appropriate types for the …
-
Jointly Predicting Predicates and Arguments in Neural Semantic Role Labeling
2018
Recent BIO-tagging-based neural semantic role labeling models are very high performing, but assume gold predicates as part of the input and cannot incorporate span-level features. We propose an endto-end approach for jointly predicting all predicates, …
-
HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case
2019 · DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)
Given a combinatorial optimisation problem, there are typically multiple ways of modelling it for presentation to an automated solver. Choosing the right combination of model and target solver can have a significant impact on the …
-
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
2020
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
-
Emergent linguistic structure in artificial neural networks trained by self-supervision
2020 · Proceedings of the National Academy of Sciences
This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is …