Researcher profile

Colin Cherry

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. What Matters Most in Morphologically Segmented SMT Models?

    2015

    Morphological segmentation is an effective strategy for addressing difficulties caused by morphological complexity. In this study, we use an English-to-Arabic test bed to determine what steps and components of a phrase-based statistical machine translation pipeline …

  2. The Unreasonable Effectiveness of Word Representations for Twitter Named Entity Recognition

    2015

    Named entity recognition (NER) systems trained on newswire perform very badly when tested on Twitter. Signals that were reliable in copy-edited text disappear almost entirely in Twitter’s informal chatter, requiring the construction of specialized models. …

  3. mSLAM: Massively multilingual joint pre-training for speech and text

    2022 · arXiv (Cornell University)

    We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unlabeled speech and text in multiple languages. mSLAM combines w2v-BERT …

  4. SMOL: Professionally translated parallel data for 115 under-represented languages

    2025 · arXiv (Cornell University)

    We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and growing) under-resourced languages (125 language pairs), including many …

  5. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  6. Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges

    2019 · arXiv (Cornell University)

    We introduce our efforts towards building a universal neural machine translation (NMT) system capable of translating between any language pair. We set a milestone towards this goal by building a single massively multilingual NMT model …