Aaron Defazio
3 papers in the PaperMetrix corpus
Papers by this author
-
Non-Uniform Stochastic Average Gradient Method for Training Conditional Random Fields
2015 · arXiv (Cornell University)
We apply stochastic average gradient (SAG) algorithms for training conditional random fields (CRFs). We describe a practical implementation that uses structure in the CRF gradient to reduce the memory requirement of this linearly-convergent stochastic gradient …
-
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
2023 · arXiv (Cornell University)
We consider the problem of estimating the learning rate in adaptive methods, such as AdaGrad and Adam. We propose Prodigy, an algorithm that provably estimates the distance to the solution $D$, which is needed to …
-
Why Gradients Rapidly Increase Near the End of Training
2025 · arXiv (Cornell University)
During long-duration Large Language Model (LLM) training runs the gradient norm increases rapidly near the end of training. In this short note, we show that this increase is due to an unintended interaction between weight …