Łukasz Kaiser
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Model-Based Reinforcement Learning for Atari
2019 · arXiv (Cornell University)
Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in …
-
Rethinking Attention with Performers
2020 · arXiv (Cornell University)
We introduce Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear (as opposed to quadratic) space and time complexity, without relying on any priors such as sparsity …
-
Multi-task Sequence to Sequence Learning
2015 · arXiv (Cornell University)
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. …
-
Sentence Compression by Deletion with LSTMs
2015
We present an LSTM approach to deletion-based sentence compression where the task is to translate a sentence into a sequence of zeros and ones, corresponding to token deletion decisions. We demonstrate that even the most …
-
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
2016 · arXiv (Cornell University)
Neural Machine Translation (NMT) is an end-to-end learning approach for automated translation, with the potential to overcome many of the weaknesses of conventional phrase-based translation systems. Unfortunately, NMT systems are known to be computationally expensive …
-
Attention Is All You Need
2025
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a …
-
Fast Decoding in Sequence Models using Discrete Latent Variables
2018 · arXiv (Cornell University)
Autoregressive sequence models based on deep neural networks, such as RNNs, Wavenet and the Transformer attain state-of-the-art results on many tasks. However, they are difficult to parallelize and are thus slow at processing long sequences. …
-
Universal Transformers
2018 · arXiv (Cornell University)
Recurrent neural networks (RNNs) sequentially process data by updating their state with each new data point, and have long been the de facto choice for sequence modeling tasks. However, their inherently sequential computation makes them …
-
Transforming machine translation: a deep learning system reaches news translation quality comparable to human professionals
2020 · Nature Communications
The quality of human translation was long thought to be unattainable for computer translation systems. In this study, we present a deep-learning system, CUBBITT, which challenges this view. In a context-aware blind evaluation by human …