Researcher profile

Yann Dauphin

8 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win

    2020 · arXiv (Cornell University)

    Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win - AAAI 2022 Poster

  2. Tied-Augment: Controlling Representation Similarity Improves Data Augmentation

    2023 · arXiv (Cornell University)

    Data augmentation methods have played an important role in the recent advance of deep learning models, and have become an indispensable component of state-of-the-art models in semi-supervised, self-supervised, and supervised training for vision. Despite incurring …

  3. A Convolutional Encoder Model for Neural Machine Translation

    2017

    The prevalent approach to neural machine translation relies on bi-directional LSTMs to encode the source sentence. We present a faster and simpler architecture based on a succession of convolutional layers. This allows to encode the …

  4. Language Modeling with Gated Convolutional Networks

    2016 · arXiv (Cornell University)

    The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a …

  5. Convolutional Sequence to Sequence Learning

    2017 · arXiv (Cornell University)

    The prevalent approach to sequence to sequence learning maps an input sequence to a variable length output sequence via recurrent neural networks. We introduce an architecture based entirely on convolutional neural networks. Compared to recurrent …

  6. Hierarchical Neural Story Generation

    2018 · arXiv (Cornell University)

    We explore story generation: creative systems that can build coherent and fluent passages of text about a topic. We collect a large dataset of 300K human-written stories paired with writing prompts from an online forum. …

  7. Pay Less Attention with Lightweight and Dynamic Convolutions

    2019 · arXiv (Cornell University)

    Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that …

  8. Language modeling with gated convolutional networks

    2017 · International Conference on Machine Learning

    The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a …