Zeyuan Allen-Zhu
3 papers in the PaperMetrix corpus
Papers by this author
-
Natasha 2: Faster Non-Convex Optimization Than SGD
2018 · Neural Information Processing Systems
(this is a theory paper) We design a stochastic algorithm to find e -approximate local minima of any smooth nonconvex function in rate O(e−3.25) , with only oracle access to stochastic gradients. The best result …
-
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
2024 · arXiv (Cornell University)
Language models have demonstrated remarkable performance in solving reasoning tasks; however, even the strongest models still occasionally make reasoning mistakes. Recently, there has been active research aimed at improving reasoning accuracy, particularly by using pretrained …
-
LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency
2021 · arXiv (Cornell University)
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes …