ملف الباحث

Zhifeng Chen

15 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Reward Augmented Maximum Likelihood for Neural Structured Prediction

    2016 · arXiv (Cornell University)

    A key problem in structured output prediction is direct optimization of the task reward function that matters for test evaluation. This paper presents a simple and computationally efficient approach to incorporate task reward into a …

  2. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

    2016 · arXiv (Cornell University)

    Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive …

  3. Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation

    2017 · Transactions of the Association for Computational Linguistics

    We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no changes to the model architecture from a standard NMT system but instead …

  4. Sequence-to-Sequence Models Can Directly Translate Foreign Speech

    2017

    We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another.The model does not explicitly transcribe the speech into text in the source language, nor does …

  5. Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

    2017 · arXiv (Cornell University)

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  6. Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

    2019 · arXiv (Cornell University)

    Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, …

  7. Gmail Smart Compose

    2019

    In this paper, we present Smart Compose, a novel system for generating interactive, real-time suggestions in Gmail that assists users in writing mails by reducing repetitive typing. In the design and deployment of such a …

  8. Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation

    2016 · arXiv (Cornell University)

    We propose a simple solution to use a single Neural Machine Translation (NMT) model to translate between multiple languages. Our solution requires no change in the model architecture from our base system but instead introduces …

  9. Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges

    2019 · arXiv (Cornell University)

    We introduce our efforts towards building a universal neural machine translation (NMT) system capable of translating between any language pair. We set a milestone towards this goal by building a single massively multilingual NMT model …

  10. Tacotron: Towards End-to-End Speech Synthesis

    2017

    A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design …

  11. Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions

    2018

    This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by …

  12. Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

    2019

    We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.Moreover, the model is able to transfer voices across languages, e.g.synthesize fluent Spanish …

  13. Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model

    2019

    We present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an intermediate text representation.The network is trained end-to-end, learning to map speech …

  14. LaMDA: Language Models for Dialog Applications

    2022 · arXiv (Cornell University)

    We present LaMDA: Language Models for Dialog Applications. LaMDA is a family of Transformer-based neural language models specialized for dialog, which have up to 137B parameters and are pre-trained on 1.56T words of public dialog …

  15. Gemini: A Family of Highly Capable Multimodal Models

    2023 · arXiv (Cornell University)

    This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging …