preprint Open access

Navigating with Graph Representations for Fast and Scalable Decoding of\n Neural Language Models

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
20
References
0
Comments
0
Paper overview

Abstract

Neural language models (NLMs) have recently gained a renewed interest by\nachieving state-of-the-art performance across many natural language processing\n(NLP) tasks. However, NLMs are very computationally demanding largely due to\nthe computational cost of the softmax layer over a large vocabulary. We observe\nthat, in decoding of many NLP tasks, only the probabilities of the top-K\nhypotheses need to be calculated preciously and K is often much smaller than\nthe vocabulary size. This paper proposes a novel softmax layer approximation\nalgorithm, called Fast Graph Decoder (FGD), which quickly identifies, for a\ngiven context, a set of K words that are most likely to occur according to a\nNLM. We demonstrate that FGD reduces the decoding time by an order of magnitude\nwhile attaining close to the full softmax baseline accuracy on neural machine\ntranslation and language modeling tasks. We also prove the theoretical\nguarantee on the softmax approximation quality.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1806.04189
OpenAlex
W2963232306
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.