GCPT: Gradient-aware Clustering Method for Efficient Post-Training Quantization in Large Neural Networks
At a glance
- Citations
- 0
- References
- 12
- Comments
- 0
Öz
Large-scale neural network models have achieved outstanding performance across diverse tasks, but often come with expensive computational costs. In this paper, we propose a gradient-aware clustering method for post-training quantization (GCPT) in order to effectively reduce the computational overhead. Our key idea is to cluster the weights of linear layers based on their gradient-aware contributions to the overall loss function. Afterwards, all weights are replaced by a small set of cluster centroids to minimize the variation of the loss function due to quantization. To further accelerate inference, the inputs associated with those weights in the same cluster are first aggregated and then the sum is multiplied with the shared centroid, thereby reducing the number of scalar multiplications. Experiments on three large-scale models demonstrate that the proposed GCPT method achieves up to 93.8% computational cost reduction, while preserving memory usage and inference accuracy, compared to other state-of-the-art methods.
Publication details
- DOI
- 10.23919/date69613.2026.11539743
- OpenAlex
- W7163511286
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.