Martin Jaggi
11 papers in the PaperMetrix corpus
Papers by this author
-
Error Feedback Fixes SignSGD and other Gradient Compression Schemes
2019 · arXiv (Cornell University)
Sign-based algorithms (e.g. signSGD) have been proposed as a biased gradient compression technique to alleviate the communication bottleneck in training large neural networks across multiple workers. We show simple convex counter-examples where signSGD does not …
-
Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
2020 · Infoscience (Ecole Polytechnique Fédérale de Lausanne)
The performance of optimizers, particularly in deep learning, depends considerably on their chosen hyperparameter configuration. The efficacy of optimizers is often studied under near-optimal problem-specific hyperparameters, and finding these settings may be prohibitively costly for …
-
Sparse Communication for Training Deep Networks
2020 · arXiv (Cornell University)
Synchronous stochastic gradient descent (SGD) is the most common method used for distributed training of deep learning models. In this algorithm, each worker shares its local gradients with others and updates the parameters using the …
-
Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
2021 · arXiv (Cornell University)
It has been experimentally observed that the efficiency of distributed training with stochastic gradient (SGD) depends decisively on the batch size and -- in asynchronous implementations -- on the gradient staleness. Especially, it has been …
-
RelaySum for Decentralized Deep Learning on Heterogeneous Data
2021 · arXiv (Cornell University)
In decentralized machine learning, workers compute model updates on their local data. Because the workers only communicate with few neighbors without central coordination, these updates propagate progressively over the network. This paradigm enables distributed training …
-
MultiModN- Multimodal, Multi-Task, Interpretable Modular Networks
2023 · arXiv (Cornell University)
Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space …
-
Intrinsic User-Centric Interpretability through Global Mixture of Experts
2024 · arXiv (Cornell University)
In human-centric settings like education or healthcare, model accuracy and model explainability are key factors for user adoption. Towards these two goals, intrinsically interpretable deep learning models have gained popularity, focusing on accurate predictions alongside …
-
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
2025 · arXiv (Cornell University)
Grokking is proposed and widely studied as an intricate phenomenon in which generalization is achieved after a long-lasting period of overfitting. In this work, we propose NeuralGrok, a novel gradient-based approach that learns an optimal …
-
TiMoE: Time-Aware Mixture of Language Experts
2025 · arXiv (Cornell University)
Large language models (LLMs) are typically trained on fixed snapshots of the web, which means that their knowledge becomes stale and their predictions risk temporal leakage: relying on information that lies in the future relative …
-
SwissCheese at SemEval-2016 Task 4: Sentiment Classification Using an Ensemble of Convolutional Neural Networks with Distant Supervision
2016
In this paper, we propose a classifier for predicting message-level sentiments of English micro-blog messages from Twitter. Our method builds upon the convolutional sentence embedding approach proposed by (Severyn and Moschitti, 2015a; Severyn and Moschitti, …
-
Unsupervised Learning of Sentence Embeddings Using Compositional n-Gram Features
2018
Matteo Pagliardini, Prakhar Gupta, Martin Jaggi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.