Researcher profile
Wang, Jinbo
1 paper in the PaperMetrix corpus
Publications
Papers by this author
-
GradPower: Powering Gradients for Faster Language Model Pre-Training
2025 · ArXiv.org
We propose GradPower, a lightweight gradient-transformation technique for accelerating language model pre-training. Given a gradient vector $g=(g_i)_i$, GradPower first applies the elementwise sign-power transformation: $φ_p(g)=({\rm sign}(g_i)|g_i|^p)_{i}$ for a fixed $p>0$, and then feeds the transformed …