preprint Open access

A Simple Way to Initialize Recurrent Networks of Rectified Linear Units

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
557
References
32
Comments
0
Paper overview

Öz

Learning long term dependencies in recurrent networks is difficult due to vanishing and exploding gradients. To overcome this difficulty, researchers have developed sophisticated optimization techniques and network architectures. In this paper, we propose a simpler solution that use recurrent neural networks composed of rectified linear units. Key to our solution is the use of the identity matrix or its scaled version to initialize the recurrent weight matrix. We find that our solution is comparable to LSTM on our four benchmarks: two toy problems involving long-range temporal structures, a large language modeling problem and a benchmark speech recognition problem.

Record transparency

Publication details

DOI
10.48550/arxiv.1504.00941
OpenAlex
W1800356822
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.