preprint وصول مفتوح

Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
64
المراجع
34
Comments
0
Paper overview

Abstract

We formulate language modeling as a matrix factorization problem, and show that the expressiveness of Softmax-based models (including the majority of neural language models) is limited by a Softmax bottleneck. Given that natural language is highly context-dependent, this further implies that in practice Softmax with distributed word embeddings does not have enough capacity to model natural language. We propose a simple and effective method to address this issue, and improve the state-of-the-art perplexities on Penn Treebank and WikiText-2 to 47.69 and 40.68 respectively. The proposed method also excels on the large-scale 1B Word dataset, outperforming the baseline by over 5.6 points in perplexity.

Record transparency

Publication details

DOI
10.48550/arxiv.1711.03953
OpenAlex
W2767321762
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.