Chris Dyer
30 papers in the PaperMetrix corpus
Papers by this author
-
Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
2015 · arXiv (Cornell University)
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
-
Frame-Semantic Role Labeling with Heterogeneous Annotations
2015
Meghana Kshirsagar, Sam Thomson, Nathan Schneider, Jaime Carbonell, Noah A. Smith, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing …
-
Segmental Recurrent Neural Networks for End-to-end Speech Recognition
2016 · arXiv (Cornell University)
We study the segmental recurrent neural network for end-to-end acoustic modelling. This model connects the segmental conditional random field (CRF) with a recurrent neural network (RNN) used for feature extraction. Compared to most previous CRF-based …
-
Learning the Curriculum with Bayesian Optimization for Task-Specific Word Representation Learning
2016
We use Bayesian optimization to learn curricula for word representation learning, optimizing performance on downstream tasks that depend on the learned representations as features. The curricula are modeled by a linear ranking function which is …
-
Generalizing and Hybridizing Count-based and Neural Language Models
2016
Language models (LMs) are statistical models that calculate probabilities over sequences of words or other discrete symbols. Currently two major paradigms for language modeling exist: count-based n-gram models, which have advantages of scalability and test-time …
-
A Continuous Relaxation of Beam Search for End-to-End Training of Neural Sequence Models
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do …
-
Learning and Evaluating General Linguistic Intelligence
2019 · arXiv (Cornell University)
We define general linguistic intelligence as the ability to reuse previously acquired knowledge about a language's lexicon, syntax, semantics, and pragmatic conventions to adapt to new tasks quickly. Using this definition, we analyze state-of-the-art natural …
-
Retrofitting Word Vectors to Semantic Lexicons
2015
Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
Improved Transition-based Parsing by Modeling Characters instead of Words with LSTMs
2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)
We present extensions to a continuousstate dependency parsing method that makes it applicable to morphologically rich languages. Starting with a highperformance transition-based parser that uses long short-term memory (LSTM) recurrent neural networks to learn representations …
-
Transition-Based Dependency Parsing with Stack Long Short-Term Memory
2015 · RECERCAT (Consorci de Serveis Universitaris de Catalunya)
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, Noah A. Smith. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: …
-
Better Hypothesis Testing for Statistical Machine Translation: Controlling for Optimizer Instability
2018 · Figshare
In statistical machine translation, a researcher seeks to determine whether some innovation (e.g., a new feature, model, or inference algorithm) improves translation quality in comparison to a baseline system. To answer this question, he runs …
-
A Simple, Fast, and Effective Reparameterization of IBM Model 2
2018 · KiltHub Repository
We present a simple log-linear reparameterization of IBM Model 2 that overcomes problems arising from Model 1’s strong assumptions and Model 2’s overparameterization. Efficient inference, likelihood evaluation, and parameter estimation algorithms are provided. Training the …
-
Character-based Neural Machine Translation
2015 · arXiv (Cornell University)
We introduce a neural machine translation model that views the input and output sentences as sequences of characters rather than words. Since word-level information provides a crucial source of bias, our input model composes representations …
-
Massively Multilingual Word Embeddings
2016 · arXiv (Cornell University)
We introduce new methods for estimating and evaluating embeddings of words in more than fifty languages in a single shared embedding space. Our estimation methods, multiCluster and multiCCA, use dictionaries and monolingual data; they do …
-
Segmental Recurrent Neural Networks
2015 · arXiv (Cornell University)
We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous subsequences …
-
Two/Too Simple Adaptations of Word2Vec for Syntax Problems
2015
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
Neural Architectures for Named Entity Recognition
2016
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
Training with Exploration Improves a Greedy Stack LSTM Parser
2016
We adapt the greedy Stack-LSTM dependency parser of Dyer et al. (2015) to support a training-with-exploration procedure using dynamic oracles(Goldberg and Nivre, 2013) instead of cross-entropy minimization. This form of training, which accounts for model …
-
Polyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation Learning
2016
Yulia Tsvetkov, Sunayana Sitaram, Manaal Faruqui, Guillaume Lample, Patrick Littell, David Mortensen, Alan W Black, Lori Levin, Chris Dyer. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: …
-
Generation from Abstract Meaning Representation using Tree Transducers
2016
Jeffrey Flanigan, Chris Dyer, Noah A. Smith, Jaime Carbonell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
Attention-based Multimodal Neural Machine Translation
2016
We present a novel neural machine translation (NMT) architecture associating visual and textual features for translation tasks with multiple modalities. Transformed global and regional visual features are concatenated with text to form attendable sequences which …
-
What Do Recurrent Neural Network Grammars Learn About Syntax?
2017
Adhiguna Kuncoro, Miguel Ballesteros, Lingpeng Kong, Chris Dyer, Graham Neubig, Noah A. Smith. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017.
-
Dynamic Integration of Background Knowledge in Neural NLU Systems
2017 · arXiv (Cornell University)
Common-sense and background knowledge is required to understand natural language, but in most neural natural language understanding (NLU) systems, this knowledge must be acquired from training corpora during learning, and then it is static at …
-
The NarrativeQA Reading Comprehension Challenge
2018 · Transactions of the Association for Computational Linguistics
Reading comprehension (RC)—in contrast to information retrieval—requires integrating information and reasoning about events, entities, and their relations across a full document. Question answering is conventionally used to assess RC ability, in both artificial agents and …
-
LSTMs Can Learn Syntax-Sensitive Dependencies Well, But Modeling Structure Makes Them Better
2018
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, Phil Blunsom. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
-
Compound Probabilistic Context-Free Grammars for Grammar Induction
2019
We study a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar. In contrast to traditional formulations which learn a single stochastic grammar, our context-free …
-
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
2017
Solving algebraic word problems requires executing a series of arithmetic operations-a program-to obtain a final answer. However, since programs can be arbitrarily complicated, inducing them directly from question-answer pairs is a formidable challenge. To make …
-
Segmental Recurrent Neural Networks
2016 · International Conference on Learning Representations
Abstract: We introduce segmental recurrent neural networks (SRNNs) which define, given an input sequence, a joint probability distribution over segmentations of the input and labelings of the segments. Representations of the input segments (i.e., contiguous …
-
Improved Part-of-Speech Tagging for Online Conversational Text with Word Clusters
2018 · Figshare
We consider the problem of part-of-speech tagging for informal, online conversational text. We systematically evaluate the use of large-scale unsupervised word clustering and new lexical features to improve tagging accuracy. With these features, our system …
-
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
2021 · arXiv (Cornell University)
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world. In this paper, we present an analysis of Transformer-based language model …