ملف الباحث
Quentin Anthony
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
RWKV: Reinventing RNNs for the Transformer Era
2023 · arXiv (Cornell University)
Transformers have revolutionized almost all natural language processing (NLP) tasks but suffer from memory and computational complexity that scales quadratically with sequence length. In contrast, recurrent neural networks (RNNs) exhibit linear scaling in memory and …
-
MCR-DL: Mix-and-Match Communication Runtime for Deep Learning
2023
In recent years, the training requirements of many state-of-the-art Deep Learning (DL) models have scaled beyond the compute and memory capabilities of a single processor, and necessitated distribution among processors. Training such massive models necessitates …