Zhewei Yao
3 papers in the PaperMetrix corpus
Papers by this author
-
Inefficiency of K-FAC for Large Batch Size Training
2019 · arXiv (Cornell University)
In stochastic optimization, using large batch sizes during training can leverage parallel resources to produce faster wall-clock training times per training epoch. However, for both training loss and testing error, recent results analyzing large batch …
-
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
2022 · arXiv (Cornell University)
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohibitive training costs. To mitigate this issue, we propose a novel random …
-
ComposeRAG: A Modular and Composable RAG for Corpus-Grounded Multi-Hop Question Answering
2025 · arXiv (Cornell University)
Retrieval-Augmented Generation (RAG) systems are increasingly diverse, yet many suffer from monolithic designs that tightly couple core functions like query reformulation, retrieval, reasoning, and verification. This limits their interpretability, systematic evaluation, and targeted improvement, especially …