Researcher profile
George E. Dahl
2 papers in the PaperMetrix corpus
Publications
Papers by this author
-
Large scale distributed neural network training through online\n distillation
2018 · arXiv (Cornell University)
Techniques such as ensembling and distillation promise model quality\nimprovements when paired with almost any base model. However, due to increased\ntest-time cost (for ensembles) and increased complexity of the training\npipeline (for distillation), these techniques are challenging …
-
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
2025 · arXiv (Cornell University)
If we want to train a neural network using any of the most popular optimization algorithms, we are immediately faced with a dilemma: how to set the various optimization and regularization hyperparameters? When computational resources …