Scaling up Deep Learning for AI using FPGAs
At a glance
- Citations
- 0
- References
- 15
- Comments
- 0
Abstract
This paper considers scaling up large Deep Learning (DL) implementations to meet the rapidly growing demand of AI applications. FPGAs offer several performance advantages for large AI implementations including SWAP, cost, throughput, latency and reprogrammability. Specifically, the focus of this paper is on large-scale implementations of dense layers networks on Multi-FPGA based systems. Dense layer networks are at the core of DL and represent a key technology area in developing and building AI systems. This paper begins with an overview and discusses dense layer networks. Due to cost and development constraints, high performance large-scale AI implementations require using FPGAs for both training and inference. The paper reviews forward and backward propagation and considers FPGA implementation for these tasks. Multi-FPGA system design is covered in the context of large-scale, high-throughput AI implementations. Finally, an example Multi-FPGA design is provided to illustrate the concepts in the paper based on the AMD Versal device.
Publication details
- DOI
- 10.1109/aero58975.2024.10521423
- OpenAlex
- W4396853366
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.