conference-paper

Accelerating the Inference of Deep Learning-based CSI Feedback on a General-purpose CPU

Research footprint

At a glance

Citations
0
References
13
Comments
0
Paper overview

Abstract

Artificial Intelligence (AI) and Machine Learning (ML) are transforming industries worldwide and the telecom industry is no exception. Nowadays, AI/ML technologies have been applied in 5G systems, mostly in network automation and non-real-time functionalities such as energy savings, load balancing, and mobility optimization. For future radio access networks, it is even more critical to investigate $\mathrm{AI} / \mathrm{ML}$ applications in real-time processing, where the AI/ML-assisted channel state information (CSI) feedback has been one of the 3GPP Rel-18 recommended use cases for AI-native radio access networks (RAN). In this paper, we investigate the feasibility and performance of a Transformer-based CSI feedback scheme on a general purpose CPU platform. We also propose acceleration schemes for this Transformer model. Our experiment results show that our proposed optimization techniques such as operator fusion, weight pre-packing, and automatic precision conversion, together with the advanced matrix extension (AMX) hardware accelerator, can significantly reduce the inference latency by $\mathbf{4 6. 8 \%}$. Our study also demonstrates that a general-purpose $\mathbf{x 6 6}$ CPU platform can support the real-time inference of AI models in 5G vRAN. In comparison to the GPU platform, we demonstrate that the CPU platform with hardware accelerator and software optimizations can achieve comparable $\mathrm{AI} / \mathrm{ML}$ inference performance; in small batch size scenarios, the CPU platform can even outperform GPU.

Record transparency

Publication details

DOI
10.1109/iccc62479.2024.10682041
OpenAlex
W4402811535
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.