Researcher profile

Chris Dyer

30 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation

    2015 · arXiv (Cornell University)

    Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.

  2. Frame-Semantic Role Labeling with Heterogeneous Annotations

    2015

    Meghana Kshirsagar, Sam Thomson, Nathan Schneider, Jaime Carbonell, Noah A. Smith, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing …

  3. Segmental Recurrent Neural Networks for End-to-end Speech Recognition

    2016 · arXiv (Cornell University)

    We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based …

  4. Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning

    2016

    We use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features. The curricula are modeled by a linear ranking function which is …

  5. Generalizing and Hybridizing Count-based and Neural Language Models

    2016

    Language models (LMs) are statistical models that calculate probabilities over sequences of words or other discrete symbols. Currently two major paradigms for language modeling exist: count-based n-gram models, which have advantages of scalability and test-time …

  6. A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models

    2018 · Proceedings of the AAAI Conference on Artificial Intelligence

    Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do …

  7. Learning and Evaluating General Linguistic Intelligence

    2019 · arXiv (Cornell University)

    We define general linguistic intelligence as the ability to reuse previously acquired knowledge about a language's lexicon, syntax, semantics, and pragmatic conventions to adapt to new tasks quickly. Using this definition, we analyze state-of-the-art natural …

  8. Retrofitting Word Vectors to Semantic Lexicons

    2015

    Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.

  9. Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs

    2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)

    We present extensions to a continuousstate dependency parsing method that makes it applicable to morphologically rich languages. Starting with a highperformance transition-based parser that uses long short-term memory (LSTM) recurrent neural networks to learn representations …

  10. Transition-Based Dependency Parsing with Stack Long Short-Term Memory

    2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)

    Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: …

  11. Better Hypothesis Testing for Statistical Machine Translation: Controlling for Optimizer Instability

    2018 · Figshare

    In statistical machine translation, a researcher seeks to determine whether some innovation (e.g., a new feature, model, or inference algorithm) improves translation quality in comparison to a baseline system. To answer this question, he runs …

  12. A Simple, Fast, and Effective Reparameterization of IBM Model 2

    2018 · KiltHub Repository

    We present a simple log-linear reparameterization of IBM Model 2 that overcomes problems arising from Model 1’s strong assumptions and Model 2’s overparameterization. Efficient inference, likelihood evaluation, and parameter estimation algorithms are provided. Training the …

  13. Character-based Neural Machine Translation

    2015 · arXiv (Cornell University)

    We introduce a neural machine translation model that views the input and output sentences as sequences of characters rather than words. Since word-level information provides a crucial source of bias, our input model composes representations …

  14. Massively Multilingual Word Embeddings

    2016 · arXiv (Cornell University)

    We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do …

  15. Segmental Recurrent Neural Networks

    2015 · arXiv (Cornell University)

    We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous subsequences …

  16. Two/Too Simple Adaptations of Word2Vec for Syntax Problems

    2015

    Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.

  17. Neural Architectures for Named Entity Recognition

    2016

    Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.

  18. Training with Exploration Improves a Greedy Stack LSTM Parser

    2016

    We adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model …

  19. Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation Learning

    2016

    Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: …

  20. Generation from Abstract Meaning Representation using Tree Transducers

    2016

    Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.

  21. Attention-based Multimodal Neural Machine Translation

    2016

    We present a novel neural machine translation (NMT) architecture associating visual and textual features for translation tasks with multiple modalities. Transformed global and regional visual features are concatenated with text to form attendable sequences which …

  22. What Do Recurrent Neural Network Grammars Learn About Syntax?

    2017

    Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.

  23. Dynamic Integration of Background Knowledge in Neural NLU Systems

    2017 · arXiv (Cornell University)

    Common-sense and background knowledge is required to understand natural language, but in most neural natural language understanding (NLU) systems, this knowledge must be acquired from training corpora during learning, and then it is static at …

  24. The NarrativeQA Reading Comprehension Challenge

    2018 · Transactions of the Association for Computational Linguistics

    Reading comprehension (RC)—in contrast to information retrieval—requires integrating information and reasoning about events, entities, and their relations across a full document. Question answering is conventionally used to assess RC ability, in both artificial agents and …

  25. LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them Better

    2018

    Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.

  26. Compound Probabilistic Context-Free Grammars for Grammar Induction

    2019

    We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar. In contrast to traditional formulations which learn a single stochastic grammar, our context-free …

  27. Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems

    2017

    Solving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer. However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge. To make …

  28. Segmental Recurrent Neural Networks

    2016 · International Conference on Learning Representations

    Abstract: We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous …

  29. Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters

    2018 · Figshare

    We consider the problem of part-of-speech tagging for informal, online conversational text. We systematically evaluate the use of large-scale unsupervised word clustering and new lexical features to improve tagging accuracy. With these features, our system …

  30. Scaling Language Models: Methods, Analysis & Insights from Training Gopher

    2021 · arXiv (Cornell University)

    Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …