conference-paper

GCPT: Gradient-aware Clustering Method for Efficient Post-Training Quantization in Large Neural Networks

Research footprint

At a glance

Citations
0
References
12
Comments
0
Paper overview

Abstract

Large-scale neural network models have achieved outstanding performance across diverse tasks, but often come with expensive computational costs. In this paper, we propose a gradient-aware clustering method for post-training quantization (GCPT) in order to effectively reduce the computational overhead. Our key idea is to cluster the weights of linear layers based on their gradient-aware contributions to the overall loss function. Afterwards, all weights are replaced by a small set of cluster centroids to minimize the variation of the loss function due to quantization. To further accelerate inference, the inputs associated with those weights in the same cluster are first aggregated and then the sum is multiplied with the shared centroid, thereby reducing the number of scalar multiplications. Experiments on three large-scale models demonstrate that the proposed GCPT method achieves up to 93.8% computational cost reduction, while preserving memory usage and inference accuracy, compared to other state-of-the-art methods.

Record transparency

Publication details

DOI
10.23919/date69613.2026.11539743
OpenAlex
W7163511286
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.