Quanquan Gu
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Stochastic Nested Variance Reduction for Nonconvex Optimization
2018 · arXiv (Cornell University)
We study finite-sum nonconvex optimization problems, where the objective function is an average of $n$ nonconvex functions. We propose a new stochastic gradient descent algorithm based on nested variance reduction. Compared with conventional stochastic variance …
-
Learning One-hidden-layer ReLU Networks via Gradient Descent
2018 · arXiv (Cornell University)
We study the problem of learning one-hidden-layer neural networks with Rectified Linear Unit (ReLU) activation function, where the inputs are sampled from standard Gaussian distribution and the outputs are generated from a noisy teacher network. …
-
On the Convergence and Robustness of Adversarial Training
2021 · arXiv (Cornell University)
Improving the robustness of deep neural networks (DNNs) to adversarial examples is an important yet challenging problem for secure deep learning. Across existing defense techniques, adversarial training with Projected Gradient Decent (PGD) is amongst the …
-
Nearly Minimax Optimal Regret for Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation
2021 · arXiv (Cornell University)
We study reinforcement learning in an infinite-horizon average-reward setting with linear function approximation, where the transition probability function of the underlying Markov Decision Process (MDP) admits a linear form over a feature mapping of the …
-
Iterative Teacher-Aware Learning
2021 · arXiv (Cornell University)
In human pedagogy, teachers and students can interact adaptively to maximize communication efficiency. The teacher adjusts her teaching method for different students, and the student, after getting familiar with the teacher's instruction mechanism, can infer …
-
Cooperative Multi-Agent Reinforcement Learning: Asynchronous Communication and Linear Function Approximation
2023 · arXiv (Cornell University)
We study multi-agent reinforcement learning in the setting of episodic Markov decision processes, where multiple agents cooperate via communication through a central server. We propose a provably efficient algorithm based on value iteration that enable …
-
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
2025 · arXiv (Cornell University)
Despite deep neural networks' powerful representation learning capabilities, theoretical understanding of how networks can simultaneously achieve meaningful feature learning and global convergence remains elusive. Existing approaches like the neural tangent kernel (NTK) are limited because …