ملف الباحث
Tokio Kajitsuka
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
2023 · arXiv (Cornell University)
Existing analyses of the expressive capacity of Transformer models have required excessively deep layers for data memorization, leading to a discrepancy with the Transformers actually used in practice. This is primarily due to the interpretation …