ملف الباحث

Xin Jiang

14 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order

    2020 · arXiv (Cornell University)

    Masked language model and autoregressive language model are two types of language models. While pretrained masked language models such as BERT overwhelm the line of natural language understanding (NLU) tasks, autoregressive language models such as …

  2. Constructing a Data Middle Platform for Promoting Renewable Energy Accommodation Capacity

    2021 · IOP Conference Series Earth and Environmental Science

    Abstract Constructing a data middle platform for renewable energy accommodation can provide essential support for the safe and efficient accommodation and the orderly development of the renewable energy industry. Based on the connotation and technical …

  3. PanGu-Bot: Efficient Generative Dialogue Pre-training from Pre-trained Language Model

    2022 · arXiv (Cornell University)

    In this paper, we introduce PanGu-Bot, a Chinese pre-trained open-domain dialogue generation model based on a large pre-trained language model (PLM) PANGU-alpha (Zeng et al.,2021). Different from other pre-trained dialogue models trained over a massive …

  4. Pre-training Language Models with Deterministic Factual Knowledge

    2022 · arXiv (Cornell University)

    Previous works show that Pre-trained Language Models (PLMs) can capture factual knowledge. However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitive to the changes of prompts when extracting factual knowledge. …

  5. Learning to Edit: Aligning LLMs with Knowledge Editing

    2024 · arXiv (Cornell University)

    Knowledge editing techniques, aiming to efficiently modify a minor proportion of knowledge in large language models (LLMs) without negatively impacting performance across other inputs, have garnered widespread attention. However, existing methods predominantly rely on memorizing …

  6. Neural Generative Question Answering

    2015 · arXiv (Cornell University)

    This paper presents an end-to-end neural network model, named Neural Generative Question Answering (GENQA), that can generate answers to simple factoid questions, based on the facts in a knowledge-base. More specifically, the model is built …

  7. Decomposable Neural Paraphrase Generation

    2019

    Paraphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level. This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence …

  8. ERNIE: Enhanced Language Representation with Informative Entities

    2019

    Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks. However, the existing …

  9. Paraphrase Generation with Deep Reinforcement Learning

    2018

    Automatic generation of paraphrases from a given sentence is an important yet challenging task in natural language processing (NLP). In this paper, we present a deep reinforcement learning approach to paraphrase generation. Specifically, we propose …

  10. Neural Generative Question Answering

    2016

    This paper presents an end-to-end neural network model, named Neural Generative Question Answering (GENQA), that can generate answers to simple factoid questions, based on the facts in a knowledge-base.More specifically, the model is built on …

  11. TinyBERT: Distilling BERT for Natural Language Understanding

    2019 · arXiv (Cornell University)

    Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resource-restricted …

  12. TinyBERT: Distilling BERT for Natural Language Understanding

    2020

    Language model pre-training, such as BERT, has significantly improved the performances of many natural language processing tasks. However, pre-trained language models are usually computationally expensive, so it is difficult to efficiently execute them on resourcerestricted …

  13. PanGu-$α$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation

    2021 · arXiv (Cornell University)

    Large-scale Pretrained Language Models (PLMs) have become the new paradigm for Natural Language Processing (NLP). PLMs with hundreds of billions parameters such as GPT-3 have demonstrated strong performances on natural language understanding and generation with …

  14. SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation

    2021 · arXiv (Cornell University)

    Code representation learning, which aims to encode the semantics of source code into distributed vectors, plays an important role in recent deep-learning-based models for code intelligence. Recently, many pre-trained language models for source code (e.g., …