ملف الباحث

Quentin Anthony

ورقتان في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. RWKV: Reinventing RNNs for the Transformer Era

    2023 · arXiv (Cornell University)

    Transformers have revolutionized almost all natural language processing (NLP) tasks but suffer from memory and computational complexity that scales quadratically with sequence length. In contrast, recurrent neural networks (RNNs) exhibit linear scaling in memory and …

  2. MCR-DL: Mix-and-Match Communication Runtime for Deep Learning

    2023

    In recent years, the training requirements of many state-of-the-art Deep Learning (DL) models have scaled beyond the compute and memory capabilities of a single processor, and necessitated distribution among processors. Training such massive models necessitates …