preprint Open access

Exploring Low-Cost Transformer Model Compression for Large-Scale Commercial Reply Suggestions

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
0
References
21
Comments
0
Paper overview

Öz

Fine-tuning pre-trained language models improves the quality of commercial reply suggestion systems, but at the cost of unsustainable training times. Popular training time reduction approaches are resource intensive, thus we explore low-cost model compression techniques like Layer Dropping and Layer Freezing. We demonstrate the efficacy of these techniques in large-data scenarios, enabling the training time reduction for a commercial email reply suggestion system by 42%, without affecting the model relevance or user engagement. We further study the robustness of these techniques to pre-trained model and dataset size ablation, and share several insights and recommendations for commercial applications.

Record transparency

Publication details

DOI
10.48550/arxiv.2111.13999
OpenAlex
W3217076640
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.