preprint
وصول مفتوح
Speeding Up Entmax
Research footprint
At a glance
- الاستشهادات
- 0
- المراجع
- 0
- Comments
- 0
Paper overview
Abstract
Softmax is the de facto standard in modern neural networks for language processing when it comes to normalizing logits. However, by producing a dense probability distribution each token in the vocabulary has a nonzero chance of being selected at each generation step, leading to a variety of reported problems in text generation. $α$-entmax of Peters et al. (2019, arXiv:1905.05702) solves this problem, but is considerably slower than softmax. In this paper, we propose an alternative to $α$-entmax, which keeps its virtuous characteristics, but is as fast as optimized softmax and achieves on par or better performance in machine translation task.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.2111.06832
- OpenAlex
- W4307869010
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.