Researcher profile

Omer Levy

14 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Improving Transformer Models by Reordering their Sublayers

    2020

    Multilayer transformer networks consist of interleaved self-attention and feedforward sublayers. Could ordering the sublayers in a different pattern lead to better performance? We generate randomly ordered transformers and train them with the language modeling objective. …

  2. What Do You Get When You Cross Beam Search with Nucleus Sampling?

    2021 · arXiv (Cornell University)

    We combine beam search with the probabilistic pruning technique of nucleus sampling to create two deterministic nucleus search algorithms for natural language generation. The first algorithm, p-exact search, locally prunes the next-token distribution and performs …

  3. Improving Distributional Similarity with Lessons Learned from Word Embeddings

    2015 · Transactions of the Association for Computational Linguistics

    Recent trends suggest that neural-network-inspired word embedding models outperform traditional count-based distributional models on word similarity and analogy detection tasks. We reveal that much of the performance gains of word embeddings are due to certain …

  4. Do Supervised Distributional Methods Really Learn Lexical Inference Relations?

    2015

    Omer Levy, Steffen Remus, Chris Biemann, Ido Dagan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.

  5. A Strong Baseline for Learning Cross-Lingual Word Embeddings from Sentence Alignments

    2017

    While cross-lingual word embeddings have been studied extensively in recent years, the qualitative differences between the different algorithms remain vague. We observe that whether or not an algorithm uses a particular feature set (sentence IDs) …

  6. Annotation Artifacts in Natural Language Inference Data

    2018

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short …

  7. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

    2018 · arXiv (Cornell University)

    For natural language understanding (NLU) technology to be maximally useful, both practically and as a scientific object of study, it must be general: it must be able to process language in a way that is …

  8. SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

    2019 · arXiv (Cornell University)

    In the last year, new models and methods for pretraining and transfer learning have driven striking performance improvements across a range of language understanding tasks. The GLUE benchmark, introduced a little over one year ago, …

  9. SpanBERT: Improving Pre-training by Representing and Predicting Spans

    2020 · Transactions of the Association for Computational Linguistics

    We present SpanBERT, a pre-training method that is designed to better represent and predict spans of text. Our approach extends BERT by (1) masking contiguous random spans, rather than random tokens, and (2) training the …

  10. Ultra-Fine Entity Typing

    2018

    We introduce a new entity typing task: given a sentence with an entity mention, the goal is to predict a set of free-form phrases (e.g. skyscraper, songwriter, or criminal) that describe appropriate types for the …

  11. Jointly Predicting Predicates and Arguments in Neural Semantic Role Labeling

    2018

    Recent BIO-tagging-based neural semantic role labeling models are very high performing, but assume gold predicates as part of the input and cannot incorporate span-level features. We propose an endto-end approach for jointly predicting all predicates, …

  12. HISTORIAE, History of Socio-Cultural Transformation as Linguistic Data Science. A Humanities Use Case

    2019 · DROPS (Schloss Dagstuhl – Leibniz Center for Informatics)

    Given a combinatorial optimisation problem, there are typically multiple ways of modelling it for presentation to an automated solver. Choosing the right combination of model and target solver can have a significant impact on the …

  13. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

    2020

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, Luke Zettlemoyer. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.

  14. Emergent linguistic structure in artificial neural networks trained by self-supervision

    2020 · Proceedings of the National Academy of Sciences

    This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is …