From Multimodal to Unimodal Attention in Transformers using Knowledge\n Distillation
At a glance
- Citations
- 5
- References
- 27
- Comments
- 0
Öz
Multimodal Deep Learning has garnered much interest, and transformers have\ntriggered novel approaches, thanks to the cross-attention mechanism. Here we\npropose an approach to deal with two key existing challenges: the high\ncomputational resource demanded and the issue of missing modalities. We\nintroduce for the first time the concept of knowledge distillation in\ntransformers to use only one modality at inference time. We report a full study\nanalyzing multiple student-teacher configurations, levels at which distillation\nis applied, and different methodologies. With the best configuration, we\nimproved the state-of-the-art accuracy by 3%, we reduced the number of\nparameters by 2.5 times and the inference time by 22%. Such\nperformance-computation tradeoff can be exploited in many applications and we\naim at opening a new research area where the deployment of complex models with\nlimited resources is demanded.\n
Publication details
- DOI
- 10.48550/arxiv.2110.08270
- OpenAlex
- W4225744324
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Oturum Açın to join the discussion.