Xiaodong Liu
22 papers in the PaperMetrix corpus
Papers by this author
-
A New Data Acquisition Client-software Model Used for Mobile Application Analysis
2016
This paper proposes a new data acquisition client-software model used for the Mobile Application Analysis. Compared with the existing software models, the model proposed in this paper can improve the reuse rate and reduce the …
-
A speculative execution strategy based on node classification and hierarchy index mechanism for heterogeneous Hadoop systems
2017
MapReduce (MR) has been widely used to process distributed large data sets. MRV2 working on Yarn, as a more advanced programing model, has gained lots of concerns. Meanwhile, speculative execution is known as an approach …
-
Unified Language Model Pre-training for Natural Language Understanding and Generation
2019 · arXiv (Cornell University)
This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, …
-
Navigating with Graph Representations for Fast and Scalable Decoding of\n Neural Language Models
2018 · arXiv (Cornell University)
Neural language models (NLMs) have recently gained a renewed interest by\nachieving state-of-the-art performance across many natural language processing\n(NLP) tasks. However, NLMs are very computationally demanding largely due to\nthe computational cost of the softmax layer over …
-
Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
2021 · ACM Transactions on Computing for Healthcare
Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A …
-
ARCH: Efficient Adversarial Regularized Training with Caching
2021
Adversarial regularization can improve model generalization in many natural language processing tasks. However, conventional approaches are computationally expensive since they need to generate a perturbation for each sample in each epoch. We propose a new …
-
Token-wise Curriculum Learning for Neural Machine Translation
2021 · arXiv (Cornell University)
Existing curriculum learning approaches to Neural Machine Translation (NMT) require sampling sufficient amounts of "easy" samples from training data at the early training stage. This is not always achievable for low-resource languages where the amount …
-
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
2022 · arXiv (Cornell University)
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …
-
Fine-Tuning Large Neural Language Models for Biomedical Natural Language Processing
2021 · arXiv (Cornell University)
Motivation: A perennial challenge for biomedical researchers and clinical practitioners is to stay abreast with the rapid growth of publications and medical notes. Natural language processing (NLP) has emerged as a promising direction for taming …
-
Parameter-efficient feature-based transfer for paraphrase identification
2022 · Natural Language Engineering
Abstract There are many types of approaches for Paraphrase Identification (PI), an NLP task of determining whether a sentence pair has equivalent semantics. Traditional approaches mainly consist of unsupervised learning and feature engineering, which are …
-
Prediction of Boiler Heat-Conducting Oil Temperature Based on Multi-Modality Fuzzy Cognitive Maps
2023
Accurate and interpretable time series prediction is of great significance in production, it can help people deal with boiler abnormalities timely and accurately, and provide guarantee for safe production. This paper proposes a time series …
-
Language Models as Inductive Reasoners
2024
Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). …
-
Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval
2015
Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, Ye-yi Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
2018 · arXiv (Cornell University)
We present a large-scale dataset, ReCoRD, for machine reading comprehension requiring commonsense reasoning. Experiments on this dataset demonstrate that the performance of state-of-the-art MRC systems fall far behind human performance. ReCoRD represents a challenge for …
-
Multi-Task Deep Neural Networks for Natural Language Understanding
2019 · arXiv (Cornell University)
In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a …
-
Cyclical Annealing Schedule: A Simple Approach to Mitigating
2019
Hao Fu, Chunyuan Li, Xiaodong Liu, Jianfeng Gao, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and …
-
Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
2019 · arXiv (Cornell University)
This paper explores the use of knowledge distillation to improve a Multi-Task Deep Neural Network (MT-DNN) (Liu et al., 2019) for learning text representations across multiple natural language understanding tasks. Although ensemble learning can improve …
-
Analysis of Points of Interests Recommended for Leisure Walk Descriptions
2024 · arXiv (Cornell University)
Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from …
-
Stochastic Answer Networks for Machine Reading Comprehension
2018
We propose a simple yet robust stochastic answer network (SAN) that simulates multi-step reasoning in machine reading comprehension. Compared to previous work such as ReasoNet which used reinforcement learning to determine the number of steps, …
-
Adversarial Training for Large Neural Language Models
2020 · arXiv (Cornell University)
Generalization and robustness are both key desiderata for designing machine learning methods. Adversarial training can enhance robustness, but past work often finds it hurts generalization. In natural language processing (NLP), pre-training large neural language models …
-
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
2020 · arXiv (Cornell University)
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture DeBERTa (Decoding-enhanced BERT with disentangled attention) that …
-
DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION
2021 · International Conference on Learning Representations
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture \textbf{DeBERTa} (\textbf{D}ecoding-\textbf{e}nhanced \textbf{BERT} with disentangled \textbf{a}ttention) that …