preprint وصول مفتوح

Transformer to CNN: Label-scarce distillation for efficient text\n classification

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

الاستشهادات
25
المراجع
0
Comments
0
Paper overview

Abstract

Significant advances have been made in Natural Language Processing (NLP)\nmodelling since the beginning of 2018. The new approaches allow for accurate\nresults, even when there is little labelled data, because these NLP models can\nbenefit from training on both task-agnostic and task-specific unlabelled data.\nHowever, these advantages come with significant size and computational costs.\nThis workshop paper outlines how our proposed convolutional student\narchitecture, having been trained by a distillation process from a large-scale\nmodel, can achieve 300x inference speedup and 39x reduction in parameter count.\nIn some cases, the student model performance surpasses its teacher on the\nstudied tasks.\n

Record transparency

Publication details

DOI
10.48550/arxiv.1909.03508
OpenAlex
W4288112596
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.