Tao Qin
17 papers in the PaperMetrix corpus
Papers by this author
-
Reinforcement Learning for Learning Rate Control
2017 · arXiv (Cornell University)
Stochastic gradient descent (SGD), which updates the model parameters by adding a local gradient times a learning rate at each step, is widely used in model training of machine learning algorithms such as neural networks. …
-
Semi-Supervised Neural Machine Translation via Marginal Distribution Estimation
2019 · IEEE/ACM Transactions on Audio Speech and Language Processing
Neural machine translation (NMT) heavily relies on parallel bilingual corpora for training. Since large-scale, high-quality parallel corpora are usually costly to collect, it is appealing to exploit monolingual corpora to improve NMT. Inspired by the …
-
Multi-Task Learning for Conversational Question Answering over a Large-Scale Knowledge Base
2019 · arXiv (Cornell University)
We consider the problem of conversational question answering over a large-scale knowledge base. To handle huge entity vocabulary of a large-scale knowledge base, recent neural semantic parsing based approaches usually decompose the task into several …
-
Task-Agnostic and Adaptive-Size BERT Compression
2021
While pre-trained language models such as BERT and RoBERTa have achieved impressive results on various natural language processing tasks, they have huge numbers of parameters and suffer from huge computational and memory costs, which make …
-
MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition
2021 · arXiv (Cornell University)
In this paper, we propose MixSpeech, a simple yet effective data augmentation method based on mixup for automatic speech recognition (ASR). MixSpeech trains an ASR model by taking a weighted combination of two different speech …
-
AQ–ABS: Anti-Quantum Attribute-based Signature for EMRs Sharing with Blockchain
2022 · 2022 IEEE Wireless Communications and Networking Conference (WCNC)
With the advancement of medical science, the implementation of Electronic Medical Records (EMRs) for enhancing the efficiency and reliability of healthcare services has become a widespread phenomenon. However, EMRs are stored in hospitals and medical …
-
RetroGraph: Retrosynthetic Planning with Graph Search
2023 · VBN Forskningsportal (Aalborg Universitet)
The data and checkpoint for "RetroGraph: Retrosynthetic Planning with Graph Search"
-
Dual Learning for Machine Translation
2016 · arXiv (Cornell University)
While neural machine translation (NMT) is making good progress in the past two years, tens of millions of bilingual sentence pairs are needed for its training. However, human labeling is very costly. To tackle this …
-
Question Answering and Question Generation as Dual Tasks
2017 · arXiv (Cornell University)
We study the problem of joint question answering (QA) and question generation (QG) in this paper. Our intuition is that QA and QG have intrinsic connections and these two tasks could improve each other. On …
-
Deliberation Networks: Sequence Generation Beyond One-Pass Decoding
2017 · Neural Information Processing Systems
The encoder-decoder framework has achieved promising progress for many sequence generation tasks, including machine translation, text summarization, dialog system, image captioning, etc. Such a framework adopts an one-pass forward process while decoding and generating a …
-
Achieving Human Parity on Automatic Chinese to English News Translation
2018 · arXiv (Cornell University)
Machine translation has made rapid advances in recent years. Millions of people are using it today in online translation systems and mobile applications in order to communicate across language barriers. The question naturally arises whether …
-
A Study of Reinforcement Learning for Neural Machine Translation
2018
Recent studies have shown that reinforcement learning (RL) is an effective approach for improving the performance of neural machine translation (NMT) system. However, due to its instability, successfully RL training is challenging, especially in real-world …
-
Layer-Wise Coordination between Encoder and Decoder for Neural Machine Translation
2018 · Neural Information Processing Systems
Neural Machine Translation (NMT) has achieved remarkable progress with the quick evolvement of model structures. In this paper, we propose the concept of layer-wise coordination for NMT, which explicitly coordinates the learning of hidden representations …
-
Multilingual Neural Machine Translation with Knowledge Distillation
2019 · arXiv (Cornell University)
Multilingual machine translation, which translates multiple languages with a single model, has attracted much attention due to its efficiency of offline training and online serving. However, traditional multilingual translation usually yields inferior accuracy compared with …
-
MASS: Masked Sequence to Sequence Pre-training for Language Generation
2019 · arXiv (Cornell University)
Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks. Inspired by the success of BERT, we propose MAsked Sequence to …
-
MPNet: Masked and Permuted Pre-training for Language Understanding
2020 · arXiv (Cornell University)
BERT adopts masked language modeling (MLM) for pre-training and is one of the most successful pre-training models. Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this …
-
BioGPT: generative pre-trained transformer for biomedical text generation and mining
2022 · Briefings in Bioinformatics
Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language …