Michael I. Jordan
5 papers in the PaperMetrix corpus
Papers by this author
-
$\ell_1$-regularized Neural Networks are Improperly Learnable in Polynomial Time
2015 · arXiv (Cornell University)
We study the improper learning of multi-layer neural networks. Suppose that the neural network to be learned has $k$ hidden layers and that the $\ell_1$-norm of the incoming weights of any neuron is bounded by …
-
Kernel feature selection via conditional covariance minimization
2017 · Neural Information Processing Systems
We propose a method for feature selection that employs kernel-based measures of independence to find a subset of covariates that is maximally predictive of the response. Building on past work in kernel dimension reduction, we …
-
Learning Stages: Phenomenon, Root Cause, Mechanism Hypothesis, and Implications.
2019 · arXiv (Cornell University)
Learning rate decay (lrDecay) is a \emph{de facto} technique for training modern neural networks. It starts with a large learning rate and then decays it multiple times. It is empirically observed to help both optimization …
-
On Learning Rates and Schrödinger Operators
2020 · arXiv (Cornell University)
The learning rate is perhaps the single most important parameter in the training of neural networks and, more broadly, in stochastic (nonconvex) optimization. Accordingly, there are numerous effective, but poorly understood, techniques for tuning the …
-
Robust Calibration with Multi-domain Temperature Scaling
2022 · arXiv (Cornell University)
Uncertainty quantification is essential for the reliable deployment of machine learning models to high-stakes application domains. Uncertainty quantification is all the more challenging when training distribution and test distribution are different, even the distribution shifts …