Minjia Zhang
3 papers in the PaperMetrix corpus
Papers by this author
-
Navigating with Graph Representations for Fast and Scalable Decoding of\n Neural Language Models
2018 · arXiv (Cornell University)
Neural language models (NLMs) have recently gained a renewed interest by\nachieving state-of-the-art performance across many natural language processing\n(NLP) tasks. However, NLMs are very computationally demanding largely due to\nthe computational cost of the softmax layer over …
-
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
2022 · arXiv (Cornell University)
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohibitive training costs. To mitigate this issue, we propose a novel random …
-
MiniKV: Pushing the Limits of LLM Inference via 2-Bit Layer-Discriminative KV Cache
2024 · arXiv (Cornell University)
How to efficiently serve LLMs in practice has become exceptionally challenging due to their prohibitive memory and computation requirements. In this study, we investigate optimizing the KV cache, whose memory footprint poses a critical bottleneck …