Yuanzhi Li
6 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Convergence Analysis of Two-layer Neural Networks with ReLU Activation
2017 · arXiv (Cornell University)
In recent years, stochastic gradient descent (SGD) based techniques has become the standard tools for training neural networks. However, formal theoretical understanding of why SGD can train neural networks in practice is largely missing. In …
-
TinyGSM: achieving >80% on GSM8k with small language models
2023 · arXiv (Cornell University)
Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving grade school math, the smallest model size so far required to …
-
Revisiting Disentanglement in Downstream Tasks: A Study on Its Necessity for Abstract Visual Reasoning
2024 · arXiv (Cornell University)
In representation learning, a disentangled representation is highly desirable as it encodes generative factors of data in a separable and compact pattern. Researchers have advocated leveraging disentangled representations to complete downstream tasks with encouraging empirical …
-
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
2024 · arXiv (Cornell University)
Language models have demonstrated remarkable performance in solving reasoning tasks; however, even the strongest models still occasionally make reasoning mistakes. Recently, there has been active research aimed at improving reasoning accuracy, particularly by using pretrained …
-
A Latent Variable Model Approach to PMI-based Word Embeddings
2016 · Transactions of the Association for Computational Linguistics
Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes …
-
LoRA Fine-Tuning of a 3B Code LLM for Algorithmic Efficiency
2021 · arXiv (Cornell University)
An important paradigm of natural language processing consists of large-scale pre-training on general domain data and adaptation to particular tasks or domains. As we pre-train larger models, full fine-tuning, which retrains all model parameters, becomes …