Researcher profile

Tie‐Yan Liu

18 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Reinforcement Learning for Learning Rate Control

    2017 · arXiv (Cornell University)

    Stochastic gradient descent (SGD), which updates the model parameters by adding a local gradient times a learning rate at each step, is widely used in model training of machine learning algorithms such as neural networks. …

  2. Learning Better Word Embedding by Asymmetric Low-Rank Projection of Knowledge Graph

    2015 · arXiv (Cornell University)

    Word embedding, which refers to low-dimensional dense vector representations of natural words, has demonstrated its power in many natural language processing tasks. However, it may suffer from the inaccurate and incomplete information contained in the …

  3. Semi-Supervised Neural Machine Translation via Marginal Distribution Estimation

    2019 · IEEE/ACM Transactions on Audio Speech and Language Processing

    Neural machine translation (NMT) heavily relies on parallel bilingual corpora for training. Since large-scale, high-quality parallel corpora are usually costly to collect, it is appealing to exploit monolingual corpora to improve NMT. Inspired by the …

  4. Asynchronous Stochastic Proximal Optimization Algorithms with Variance Reduction

    2017 · Proceedings of the AAAI Conference on Artificial Intelligence

    Regularized empirical risk minimization (R-ERM) is an important branch of machine learning, since it constrains the capacity of the hypothesis space and guarantees the generalization ability of the learning algorithm. Two classic proximal optimization algorithms, …

  5. Task-Agnostic and Adaptive-Size BERT Compression

    2021

    While pre-trained language models such as BERT and RoBERTa have achieved impressive results on various natural language processing tasks, they have huge numbers of parameters and suffer from huge computational and memory costs, which make …

  6. Indiscriminate Poisoning Attacks Are Shortcuts.

    2021 · arXiv (Cornell University)

    Indiscriminate data poisoning attacks, which add imperceptible perturbations to training data to maximize the test error of trained models, have become a trendy topic because they are thought to be capable of preventing unauthorized use …

  7. METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals

    2022 · arXiv (Cornell University)

    We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …

  8. Bag-of-Entities Representation for Ranking

    2016

    This paper presents a new bag-of-entities representation for document ranking, with the help of modern knowledge bases and automatic entity linking. Our system represents query and documents by bag-of-entities vectors constructed from their entity annotations, …

  9. Dual Learning for Machine Translation

    2016 · arXiv (Cornell University)

    While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this …

  10. Deliberation Networks: Sequence Generation Beyond One-Pass Decoding

    2017 · Neural Information Processing Systems

    The encoder-decoder framework has achieved promising progress for many sequence generation tasks, including machine translation, text summarization, dialog system, image captioning, etc. Such a framework adopts an one-pass forward process while decoding and generating a …

  11. Achieving Human Parity on Automatic Chinese to English News Translation

    2018 · arXiv (Cornell University)

    Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether …

  12. A Study of Reinforcement Learning for Neural Machine Translation

    2018

    Recent studies have shown that reinforcement learning (RL) is an effective approach for improving the performance of neural machine translation (NMT) system. However, due to its instability, successfully RL training is challenging, especially in real-world …

  13. Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation

    2018 · Neural Information Processing Systems

    Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations …

  14. Multilingual Neural Machine Translation with Knowledge Distillation

    2019 · arXiv (Cornell University)

    Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with …

  15. MASS: Masked Sequence to Sequence Pre-training for Language Generation

    2019 · arXiv (Cornell University)

    Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to …

  16. MPNet: Masked and Permuted Pre-training for Language Understanding

    2020 · arXiv (Cornell University)

    BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this …

  17. Rethinking Positional Encoding in Language Pre-training

    2020 · arXiv (Cornell University)

    In this work, we investigate the positional encoding methods used in language pre-training (e.g., BERT) and identify several problems in the existing formulations. First, we show that in the absolute positional encoding, the addition operation …

  18. BioGPT: generative pre-trained transformer for biomedical text generation and mining

    2022 · Briefings in Bioinformatics

    Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language …