Researcher profile

Michael Collins

10 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency

    2018

    Noise Contrastive Estimation (NCE) is a powerful parameter estimation method for loglinear models, which avoids calculation of the partition function or its derivatives at each training step, a computationally demanding step in many cases. It …

  2. A Biologically Plausible Parser

    2021 · arXiv (Cornell University)

    We describe a parser of English effectuated by biologically plausible neurons and synapses, and implemented through the Assembly Calculus, a recently proposed computational framework for cognitive function. We demonstrate that this device is capable of …

  3. Structured Training for Neural Network Transition-Based Parsing

    2015

    David Weiss, Chris Alberti, Michael Collins, Slav Petrov. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.

  4. Transforming Dependency Structures to Logical Forms for Semantic Parsing

    2016 · Transactions of the Association for Computational Linguistics

    The strongly typed syntax of grammar formalisms such as CCG, TAG, LFG and HPSG offers a synchronous framework for deriving syntactic structures and semantic logical forms. In contrast—partly due to the lack of a strong …

  5. Globally Normalized Transition-Based Neural Networks

    2016

    Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, Michael Collins. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.

  6. Unsupervised Part-Of-Speech Tagging with Anchor Hidden Markov Models

    2016 · Transactions of the Association for Computational Linguistics

    We tackle unsupervised part-of-speech (POS) tagging by learning hidden Markov models (HMMs) that are particularly well-suited for the problem. These HMMs, which we call anchor HMMs, assume that each tag is associated with at least …

  7. A BERT Baseline for the Natural Questions

    2019 · arXiv (Cornell University)

    This technical note describes a new baseline for the Natural Questions. Our model is based on BERT and reduces the gap between the model F1 scores reported in the original dataset paper and the human …

  8. Natural Questions: A Benchmark for Question Answering Research

    2019 · Transactions of the Association for Computational Linguistics

    We present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia …

  9. Synthetic QA Corpora Generation with Roundtrip Consistency

    2019

    We introduce a novel method of generating synthetic question answering corpora by combining models of question generation and answer extraction, and by filtering the results to ensure roundtrip consistency. By pretraining on the resulting corpora …

  10. T<scp>y</scp>D<scp>i</scp> QA: A Benchmark for Information-Seeking Question Answering in <i>Ty</i>pologically <i>Di</i>verse Languages

    2020 · Transactions of the Association for Computational Linguistics

    Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA—a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse with regard …