Jianfeng Lu
4 papers in the PaperMetrix corpus
Papers by this author
-
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
2023 · arXiv (Cornell University)
We study the convergence of stochastic gradient descent (SGD) for non-convex objective functions. We establish the local convergence with positive probability under the local Łojasiewicz condition introduced by Chatterjee in \cite{chatterjee2022convergence} and an additional local …
-
Accelerate Langevin Sampling with Birth-Death Process and Exploration Component
2023 · arXiv (Cornell University)
Sampling a probability distribution with known likelihood is a fundamental task in computational science and engineering. Aiming at multimodality, we propose a new sampling method that takes advantage of both birth-death process and exploration component. …
-
MASS: Masked Sequence to Sequence Pre-training for Language Generation
2019 · arXiv (Cornell University)
Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to …
-
MPNet: Masked and Permuted Pre-training for Language Understanding
2020 · arXiv (Cornell University)
BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this …