Tie‐Yan Liu
18 papers in the PaperMetrix corpus
Papers by this author
-
Reinforcement Learning for Learning Rate Control
2017 · arXiv (Cornell University)
Stochastic gradient descent (SGD), which updates the model parameters by adding a local gradient times a learning rate at each step, is widely used in model training of machine learning algorithms such as neural networks. …
-
Learning Better Word Embedding by Asymmetric Low-Rank Projection of Knowledge Graph
2015 · arXiv (Cornell University)
Word embedding, which refers to low-dimensional dense vector representations of natural words, has demonstrated its power in many natural language processing tasks. However, it may suffer from the inaccurate and incomplete information contained in the …
-
Semi-Supervised Neural Machine Translation via Marginal Distribution Estimation
2019 · IEEE/ACM Transactions on Audio Speech and Language Processing
Neural machine translation (NMT) heavily relies on parallel bilingual corpora for training. Since large-scale, high-quality parallel corpora are usually costly to collect, it is appealing to exploit monolingual corpora to improve NMT. Inspired by the …
-
Asynchronous Stochastic Proximal Optimization Algorithms with Variance Reduction
2017 · Proceedings of the AAAI Conference on Artificial Intelligence
Regularized empirical risk minimization (R-ERM) is an important branch of machine learning, since it constrains the capacity of the hypothesis space and guarantees the generalization ability of the learning algorithm. Two classic proximal optimization algorithms, …
-
Task-Agnostic and Adaptive-Size BERT Compression
2021
While pre-trained language models such as BERT and RoBERTa have achieved impressive results on various natural language processing tasks, they have huge numbers of parameters and suffer from huge computational and memory costs, which make …
-
Indiscriminate Poisoning Attacks Are Shortcuts.
2021 · arXiv (Cornell University)
Indiscriminate data poisoning attacks, which add imperceptible perturbations to training data to maximize the test error of trained models, have become a trendy topic because they are thought to be capable of preventing unauthorized use …
-
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
2022 · arXiv (Cornell University)
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …
-
Bag-of-Entities Representation for Ranking
2016
This paper presents a new bag-of-entities representation for document ranking, with the help of modern knowledge bases and automatic entity linking. Our system represents query and documents by bag-of-entities vectors constructed from their entity annotations, …
-
Dual Learning for Machine Translation
2016 · arXiv (Cornell University)
While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this …
-
Deliberation Networks: Sequence Generation Beyond One-Pass Decoding
2017 · Neural Information Processing Systems
The encoder-decoder framework has achieved promising progress for many sequence generation tasks, including machine translation, text summarization, dialog system, image captioning, etc. Such a framework adopts an one-pass forward process while decoding and generating a …
-
Achieving Human Parity on Automatic Chinese to English News Translation
2018 · arXiv (Cornell University)
Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether …
-
A Study of Reinforcement Learning for Neural Machine Translation
2018
Recent studies have shown that reinforcement learning (RL) is an effective approach for improving the performance of neural machine translation (NMT) system. However, due to its instability, successfully RL training is challenging, especially in real-world …
-
Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation
2018 · Neural Information Processing Systems
Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations …
-
Multilingual Neural Machine Translation with Knowledge Distillation
2019 · arXiv (Cornell University)
Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with …
-
MASS: Masked Sequence to Sequence Pre-training for Language Generation
2019 · arXiv (Cornell University)
Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to …
-
MPNet: Masked and Permuted Pre-training for Language Understanding
2020 · arXiv (Cornell University)
BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this …
-
Rethinking Positional Encoding in Language Pre-training
2020 · arXiv (Cornell University)
In this work, we investigate the positional encoding methods used in language pre-training (e.g., BERT) and identify several problems in the existing formulations. First, we show that in the absolute positional encoding, the addition operation …
-
BioGPT: generative pre-trained transformer for biomedical text generation and mining
2022 · Briefings in Bioinformatics
Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language …