Researcher profile

Yong Cheng

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Abnormal Client Behavior Detection in Federated Learning

    2019 · arXiv (Cornell University)

    In federated learning systems, clients are autonomous in that their behaviors are not fully governed by the server. Consequently, a client may intentionally or unintentionally deviate from the prescribed course of federated model training, resulting …

  2. mSLAM: Massively multilingual joint pre-training for speech and text

    2022 · arXiv (Cornell University)

    We present mSLAM, a multilingual Speech and LAnguage Model that learns cross-lingual cross-modal representations of speech and text by pre-training jointly on large amounts of unlabeled speech and text in multiple languages. mSLAM combines w2v-BERT …

  3. Semi-Supervised Learning for Neural Machine Translation

    2016 · arXiv (Cornell University)

    While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource …

  4. THUMT: An Open Source Toolkit for Neural Machine Translation

    2017 · arXiv (Cornell University)

    This paper introduces THUMT, an open-source toolkit for neural machine translation (NMT) developed by the Natural Language Processing Group at Tsinghua University. THUMT implements the standard attention-based encoder-decoder framework on top of Theano and supports …

  5. Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach

    2019

    While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors. In this work, we propose a contrastive learning approach to reducing word omission errors in NMT. The …

  6. Minimum Risk Training for Neural Machine Translation

    2016

    We propose minimum risk training for end-to-end neural machine translation. Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily …

  7. Gemini: A Family of Highly Capable Multimodal Models

    2023 · arXiv (Cornell University)

    This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging …