Di He
11 papers in the PaperMetrix corpus
Papers by this author
-
Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks
2018 · arXiv (Cornell University)
Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions …
-
MACER: Attack-free and Scalable Robust Training via Maximizing Certified Radius
2020 · arXiv (Cornell University)
Adversarial training is one of the most popular ways to learn robust models but is usually attack-dependent and time costly. In this paper, we propose the MACER algorithm, which learns robust models without using adversarial …
-
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
2022 · arXiv (Cornell University)
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …
-
Hebbian Learning based Orthogonal Projection for Continual Learning of Spiking Neural Networks
2024 · arXiv (Cornell University)
Neuromorphic computing with spiking neural networks is promising for energy-efficient artificial intelligence (AI) applications. However, different from humans who continually learn different tasks in a lifetime, neural network models suffer from catastrophic forgetting. How could …
-
REST: Retrieval-Based Speculative Decoding
2024
Zhenyu He, Zexuan Zhong, Tianle Cai, Jason Lee, Di He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
-
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning
2025 · arXiv (Cornell University)
Chain-of-Thought (CoT) prompting has emerged as a powerful technique for enhancing language model's reasoning capabilities. However, generating long and correct CoT trajectories is challenging. Recent studies have demonstrated that Looped Transformers possess remarkable length generalization …
-
Distributed Interference-Aware Power Optimization for Multi-Task Over-the-Air Federated Learning
2025 · Telecom
Over-the-air federated learning (Air-FL) has emerged as a promising paradigm that integrates communication and learning, which offers significant potential to enhance model training efficiency and optimize communication resource utilization. This paper addresses the challenge of …
-
Dual Learning for Machine Translation
2016 · arXiv (Cornell University)
While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this …
-
Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation
2018 · Neural Information Processing Systems
Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations …
-
Multilingual Neural Machine Translation with Knowledge Distillation
2019 · arXiv (Cornell University)
Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with …
-
Rethinking Positional Encoding in Language Pre-training
2020 · arXiv (Cornell University)
In this work, we investigate the positional encoding methods used in language pre-training (e.g., BERT) and identify several problems in the existing formulations. First, we show that in the absolute positional encoding, the addition operation …