preprint
Open access
Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
Research footprint
At a glance
- Citations
- 16
- References
- 12
- Comments
- 0
Paper overview
Öz
We investigate the parameter-space geometry of recurrent neural networks (RNNs), and develop an adaptation of path-SGD optimization method, attuned to this geometry, that can learn plain RNNs with ReLU activations. On several datasets that require capturing long-term dependency structure, we show that path-SGD can significantly improve trainability of ReLU RNNs compared to RNNs trained with SGD, even with various recently suggested initialization schemes.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1605.07154
- OpenAlex
- W2401137308
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Oturum Açın to join the discussion.