article

Large-scale training of deep neural networks

  • Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign)
  • University of Illinois System
Research footprint

At a glance

الاستشهادات
0
المراجع
0
Comments
0
Paper overview

Abstract

Accelerating and scaling the training of deep neural networks (DNNs) is critical to keep up with growing datasets, reduce training times, and enable training on memory-constrained problems where parallelism is necessary. In this thesis, I present a set of techniques that can leverage large high-performance computing systems for fast training of DNNs. I first introduce a suite of algorithms to exploit additional parallelism in convolutional layers when training, expanding beyond the standard sample-wise data-parallel approach to include spatial parallelism and channel and filter parallelism. Next, I present optimizations to communication frameworks to reduce communication overheads at large scales. Finally, I discuss communication quantization, which can directly reduce communication volumes. In concert, these methods allow rapid training and enable training on problems that were previously infeasible.

Record transparency

Publication details

OpenAlex
W3196474674
Document type
article
Language
EN
Source
Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.