ملف الباحث

Han Zhao

ورقة واحدة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers

    2024

    The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks.However, the effectiveness of large language models are reliant on an exponentially increasing number of parameters.The overwhelming computation complexity …