Researcher profile

Di He

11 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks

    2018 · arXiv (Cornell University)

    Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions …

  2. MACER: Attack-free and Scalable Robust Training via Maximizing Certified Radius

    2020 · arXiv (Cornell University)

    Adversarial training is one of the most popular ways to learn robust models but is usually attack-dependent and time costly. In this paper, we propose the MACER algorithm, which learns robust models without using adversarial …

  3. METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals

    2022 · arXiv (Cornell University)

    We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …

  4. Hebbian Learning based Orthogonal Projection for Continual Learning of Spiking Neural Networks

    2024 · arXiv (Cornell University)

    Neuromorphic computing with spiking neural networks is promising for energy-efficient artificial intelligence (AI) applications. However, different from humans who continually learn different tasks in a lifetime, neural network models suffer from catastrophic forgetting. How could …

  5. REST: Retrieval-Based Speculative Decoding

    2024

    Zhenyu He, Zexuan Zhong, Tianle Cai, Jason Lee, Di He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.

  6. Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning

    2025 · arXiv (Cornell University)

    Chain-of-Thought (CoT) prompting has emerged as a powerful technique for enhancing language model's reasoning capabilities. However, generating long and correct CoT trajectories is challenging. Recent studies have demonstrated that Looped Transformers possess remarkable length generalization …

  7. Distributed Interference-Aware Power Optimization for Multi-Task Over-the-Air Federated Learning

    2025 · Telecom

    Over-the-air federated learning (Air-FL) has emerged as a promising paradigm that integrates communication and learning, which offers significant potential to enhance model training efficiency and optimize communication resource utilization. This paper addresses the challenge of …

  8. Dual Learning for Machine Translation

    2016 · arXiv (Cornell University)

    While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this …

  9. Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation

    2018 · Neural Information Processing Systems

    Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations …

  10. Multilingual Neural Machine Translation with Knowledge Distillation

    2019 · arXiv (Cornell University)

    Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with …

  11. Rethinking Positional Encoding in Language Pre-training

    2020 · arXiv (Cornell University)

    In this work, we investigate the positional encoding methods used in language pre-training (e.g., BERT) and identify several problems in the existing formulations. First, we show that in the absolute positional encoding, the addition operation …