preprint
وصول مفتوح
Transformer to CNN: Label-scarce distillation for efficient text\n classification
Research footprint
At a glance
- الاستشهادات
- 25
- المراجع
- 0
- Comments
- 0
Paper overview
Abstract
Significant advances have been made in Natural Language Processing (NLP)\nmodelling since the beginning of 2018. The new approaches allow for accurate\nresults, even when there is little labelled data, because these NLP models can\nbenefit from training on both task-agnostic and task-specific unlabelled data.\nHowever, these advantages come with significant size and computational costs.\nThis workshop paper outlines how our proposed convolutional student\narchitecture, having been trained by a distillation process from a large-scale\nmodel, can achieve 300x inference speedup and 39x reduction in parameter count.\nIn some cases, the student model performance surpasses its teacher on the\nstudied tasks.\n
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1909.03508
- OpenAlex
- W4288112596
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.