conference-paper

An Empirical Study of Issues in Large Language Model Training Systems

Research footprint

At a glance

Citations
1
References
29
Comments
0
Paper overview

Abstract

Large language models (LLMs) have gained significant traction in recent years, driving advancements in various applications. The training and evaluation of these models depend heavily on specialized LLM training systems, which are deployed across numerous GPUs, partition LLMs, and process large datasets. However, issues in LLM training systems can lead to program crashes or unexpected behavior, reducing development productivity and wasting valuable resources such as GPUs and storage.

Record transparency

Publication details

DOI
10.1145/3696630.3728538
OpenAlex
W4413268030
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.