Researcher profile

Xiaodong Liu

22 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. A New Data Acquisition Client-software Model Used for Mobile Application Analysis

    2016

    This paper proposes a new data acquisition client-software model used for the Mobile Application Analysis. Compared with the existing software models, the model proposed in this paper can improve the reuse rate and reduce the …

  2. A speculative execution strategy based on node classification and hierarchy index mechanism for heterogeneous Hadoop systems

    2017

    MapReduce (MR) has been widely used to process distributed large data sets. MRV2 working on Yarn, as a more advanced programing model, has gained lots of concerns. Meanwhile, speculative execution is known as an approach …

  3. Unified Language Model Pre-training for Natural Language Understanding and Generation

    2019 · arXiv (Cornell University)

    This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, …

  4. Navigating with Graph Representations for Fast and Scalable Decoding of\n Neural Language Models

    2018 · arXiv (Cornell University)

    Neural language models (NLMs) have recently gained a renewed interest by\nachieving state-of-the-art performance across many natural language processing\n(NLP) tasks. However, NLMs are very computationally demanding largely due to\nthe computational cost of the softmax layer over …

  5. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

    2021 · ACM Transactions on Computing for Healthcare

    Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A …

  6. ARCH: Efficient Adversarial Regularized Training with Caching

    2021

    Adversarial regularization can improve model generalization in many natural language processing tasks. However, conventional approaches are computationally expensive since they need to generate a perturbation for each sample in each epoch. We propose a new …

  7. Token-wise Curriculum Learning for Neural Machine Translation

    2021 · arXiv (Cornell University)

    Existing curriculum learning approaches to Neural Machine Translation (NMT) require sampling sufficient amounts of "easy" samples from training data at the early training stage. This is not always achievable for low-resource languages where the amount …

  8. METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals

    2022 · arXiv (Cornell University)

    We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …

  9. Fine-Tuning Large Neural Language Models for Biomedical Natural Language Processing

    2021 · arXiv (Cornell University)

    Motivation: A perennial challenge for biomedical researchers and clinical practitioners is to stay abreast with the rapid growth of publications and medical notes. Natural language processing (NLP) has emerged as a promising direction for taming …

  10. Parameter-efficient feature-based transfer for paraphrase identification

    2022 · Natural Language Engineering

    Abstract There are many types of approaches for Paraphrase Identification (PI), an NLP task of determining whether a sentence pair has equivalent semantics. Traditional approaches mainly consist of unsupervised learning and feature engineering, which are …

  11. Prediction of Boiler Heat-Conducting Oil Temperature Based on Multi-Modality Fuzzy Cognitive Maps

    2023

    Accurate and interpretable time series prediction is of great significance in production, it can help people deal with boiler abnormalities timely and accurately, and provide guarantee for safe production. This paper proposes a time series …

  12. Language Models as Inductive Reasoners

    2024

    Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). …

  13. Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval

    2015

    Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, Ye-yi Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.

  14. ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension

    2018 · arXiv (Cornell University)

    We present a large-scale dataset, ReCoRD, for machine reading comprehension requiring commonsense reasoning. Experiments on this dataset demonstrate that the performance of state-of-the-art MRC systems fall far behind human performance. ReCoRD represents a challenge for …

  15. Multi-Task Deep Neural Networks for Natural Language Understanding

    2019 · arXiv (Cornell University)

    In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a …

  16. Cyclical Annealing Schedule: A Simple Approach to Mitigating

    2019

    Hao Fu, Chunyuan Li, Xiaodong Liu, Jianfeng Gao, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and …

  17. Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding

    2019 · arXiv (Cornell University)

    This paper explores the use of knowledge distillation to improve a Multi-Task Deep Neural Network (MT-DNN) (Liu et al., 2019) for learning text representations across multiple natural language understanding tasks. Although ensemble learning can improve …

  18. Analysis of Points of Interests Recommended for Leisure Walk Descriptions

    2024 · arXiv (Cornell University)

    Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from …

  19. Stochastic Answer Networks for Machine Reading Comprehension

    2018

    We propose a simple yet robust stochastic answer network (SAN) that simulates multi-step reasoning in machine reading comprehension. Compared to previous work such as ReasoNet which used reinforcement learning to determine the number of steps, …

  20. Adversarial Training for Large Neural Language Models

    2020 · arXiv (Cornell University)

    Generalization and robustness are both key desiderata for designing machine learning methods. Adversarial training can enhance robustness, but past work often finds it hurts generalization. In natural language processing (NLP), pre-training large neural language models …

  21. DeBERTa: Decoding-enhanced BERT with Disentangled Attention

    2020 · arXiv (Cornell University)

    Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture DeBERTa (Decoding-enhanced BERT with disentangled attention) that …

  22. DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION

    2021 · International Conference on Learning Representations

    Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture \textbf{DeBERTa} (\textbf{D}ecoding-\textbf{e}nhanced \textbf{BERT} with disentangled \textbf{a}ttention) that …