Jianfeng Gao
46 papers in the PaperMetrix corpus
Papers by this author
-
Modeling Large-Scale Structured Relationships with Shared Memory for Knowledge Base Completion
2017
Recent studies on knowledge base completion, the task of recovering missing relationships based on recorded relations, demonstrate the importance of learning embeddings from multi-step relations. However, due to the size of knowledge bases, learning multi-step …
-
ReasoNet: Learning to Stop Reading in Machine Comprehension
2016 · arXiv (Cornell University)
Teaching a computer to read and answer general questions pertaining to a document is a challenging yet unsolved problem. In this paper, we describe a novel neural network architecture called the Reasoning Network (ReasoNet) for …
-
Multi-Task Learning of Speaker-Role-Based Neural Conversation Models
2017 · International Joint Conference on Natural Language Processing
Building a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training neural …
-
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning
2018
Training a task-completion dialogue agent via reinforcement learning (RL) is costly because it requires many interactions with real users. One common alternative is to use a user simulator. However, a user simulator usually lacks the …
-
Unified Language Model Pre-training for Natural Language Understanding and Generation
2019 · arXiv (Cornell University)
This paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks. The model is pre-trained using three types of language modeling tasks: unidirectional, bidirectional, …
-
deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
2015 · arXiv (Cornell University)
We introduce Discriminative BLEU (deltaBLEU), a novel metric for intrinsic evaluation of generated text in tasks that admit a diverse range of possible outputs. Reference strings are scored for quality by human raters on a …
-
End-to-End Joint Learning of Natural Language Understanding and Dialogue Manager
2016 · arXiv (Cornell University)
Natural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding …
-
Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
2017
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, Li Deng. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
-
Navigating with Graph Representations for Fast and Scalable Decoding of\n Neural Language Models
2018 · arXiv (Cornell University)
Neural language models (NLMs) have recently gained a renewed interest by\nachieving state-of-the-art performance across many natural language processing\n(NLP) tasks. However, NLMs are very computationally demanding largely due to\nthe computational cost of the softmax layer over …
-
Deep Reinforcement Learning with a Natural Language Action Space
2016
Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, Mari Ostendorf. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
-
Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving
2019 · arXiv (Cornell University)
We incorporate Tensor-Product Representations within the Transformer in order to better support the explicit representation of relation structure. Our Tensor-Product Transformer (TP-Transformer) sets a new state of the art on the recently-introduced Mathematics Dataset containing …
-
Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
2021 · ACM Transactions on Computing for Healthcare
Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A …
-
ARCH: Efficient Adversarial Regularized Training with Caching
2021
Adversarial regularization can improve model generalization in many natural language processing tasks. However, conventional approaches are computationally expensive since they need to generate a perturbation for each sample in each epoch. We propose a new …
-
Token-wise Curriculum Learning for Neural Machine Translation
2021 · arXiv (Cornell University)
Existing curriculum learning approaches to Neural Machine Translation (NMT) require sampling sufficient amounts of "easy" samples from training data at the early training stage. This is not always achievable for low-resource languages where the amount …
-
METRO: Efficient Denoising Pretraining of Large Scale Autoencoding Language Models with Model Generated Signals
2022 · arXiv (Cornell University)
We present an efficient method of pretraining large-scale autoencoding language models using training signals generated by an auxiliary model. Originated in ELECTRA, this training strategy has demonstrated sample-efficiency to pretrain models at the scale of …
-
Fine-Tuning Large Neural Language Models for Biomedical Natural Language Processing
2021 · arXiv (Cornell University)
Motivation: A perennial challenge for biomedical researchers and clinical practitioners is to stay abreast with the rapid growth of publications and medical notes. Natural language processing (NLP) has emerged as a promising direction for taming …
-
Language Models as Inductive Reasoners
2024
Zonglin Yang, Li Dong, Xinya Du, Hao Cheng, Erik Cambria, Xiaodong Liu, Jianfeng Gao, Furu Wei. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). …
-
A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
2015
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human …
-
A Diversity-Promoting Objective Function for Neural Conversation Models
2015 · arXiv (Cornell University)
Sequence-to-sequence neural network models for generation of conversational responses tend to generate safe, commonplace responses (e.g., "I don't know") regardless of the input. We suggest that the traditional objective function, i.e., the likelihood of output …
-
Semantic Parsing via Staged Query Graph Generation: Question Answering with Knowledge Base
2015
Wen-tau Yih, Ming-Wei Chang, Xiaodong He, Jianfeng Gao. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
-
Representation Learning Using Multi-Task Deep Neural Networks for Semantic Classification and Information Retrieval
2015
Xiaodong Liu, Jianfeng Gao, Xiaodong He, Li Deng, Kevin Duh, Ye-yi Wang. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
-
A Persona-Based Neural Conversation Model
2016 · arXiv (Cornell University)
We present persona-based models for handling the issue of speaker consistency in neural response generation. A speaker model encodes personas in distributed embeddings that capture individual characteristics such as background information and speaking style. A …
-
End-to-End Memory Networks with Knowledge Carryover for Multi-Turn Spoken Language Understanding
2016
Spoken language understanding (SLU) is a core component of a spoken dialogue system. In the traditional architecture of dialogue systems, the SLU component treats each utterance independent of each other, and then the following components …
-
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.
2016 · Neural Information Processing Systems
This paper presents our recent work on the design and development of a new, large scale dataset, which we name MS MARCO, for MAchine Reading COmprehension. This new dataset is aimed to overcome a number …
-
A Knowledge-Grounded Neural Conversation Model
2018 · Proceedings of the AAAI Conference on Artificial Intelligence
Neural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more …
-
Neural Approaches to Conversational AI
2018
This tutorial surveys neural approaches to conversational AI that were developed in the last few years. We group conversational systems into three categories: (1) question answering agents, (2) task-oriented dialogue agents, and (3) social bots. …
-
ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
2018 · arXiv (Cornell University)
We present a large-scale dataset, ReCoRD, for machine reading comprehension requiring commonsense reasoning. Experiments on this dataset demonstrate that the performance of state-of-the-art MRC systems fall far behind human performance. ReCoRD represents a challenge for …
-
Multi-Task Deep Neural Networks for Natural Language Understanding
2019 · arXiv (Cornell University)
In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a …
-
Jointly Optimizing Diversity and Relevance in Neural Response Generation
2019
Xiang Gao, Sungjin Lee, Yizhe Zhang, Chris Brockett, Michel Galley, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 …
-
Cyclical Annealing Schedule: A Simple Approach to Mitigating
2019
Hao Fu, Chunyuan Li, Xiaodong Liu, Jianfeng Gao, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and …
-
Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
2019 · arXiv (Cornell University)
This paper explores the use of knowledge distillation to improve a Multi-Task Deep Neural Network (MT-DNN) (Liu et al., 2019) for learning text representations across multiple natural language understanding tasks. Although ensemble learning can improve …
-
Analysis of Points of Interests Recommended for Leisure Walk Descriptions
2024 · arXiv (Cornell University)
Data for Sub-Task 1 of the Advertisement in Retrieval-Augmented Generation task at Touché 2025. The dataset contains segments retrieved from the segmented version of MS MARCO V2.1. The queries used in retrieval are taken from …
-
A Nested Attention Neural Hybrid Model for Grammatical Error Correction
2017
Jianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
-
A Diversity-Promoting Objective Function for Neural Conversation Models
2016
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
-
Stochastic Answer Networks for Machine Reading Comprehension
2018
We propose a simple yet robust stochastic answer network (SAN) that simulates multi-step reasoning in machine reading comprehension. Compared to previous work such as ReasoNet which used reinforcement learning to determine the number of steps, …
-
DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation
2020
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, Bill Dolan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. 2020.
-
PIQA: Reasoning about Physical Commonsense in Natural Language
2020
To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to today's natural language understanding systems. While recent pretrained models …
-
Adversarial Training for Large Neural Language Models
2020 · arXiv (Cornell University)
Generalization and robustness are both key desiderata for designing machine learning methods. Adversarial training can enhance robustness, but past work often finds it hurts generalization. In natural language processing (NLP), pre-training large neural language models …
-
A Controllable Model of Grounded Response Generation
2021 · Proceedings of the AAAI Conference on Artificial Intelligence
Current end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, …
-
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
2020 · arXiv (Cornell University)
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture DeBERTa (Decoding-enhanced BERT with disentangled attention) that …
-
MIND: A Large-scale Dataset for News Recommendation
2020
Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, Ming Zhou. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
-
Evaluation of Text Generation: A Survey
2020 · arXiv (Cornell University)
The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years. We group NLG evaluation methods into three categories: (1) human-centric evaluation metrics, (2) automatic metrics …
-
DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation
2019 · arXiv (Cornell University)
We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer). Trained on 147M conversation-like exchanges extracted from Reddit comment chains over a period spanning from 2005 through 2017, DialoGPT extends …
-
Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space
2020
When trained effectively, the Variational Autoencoder (VAE) In this paper, we propose the first large-scale language VAE model OPTIMUS 1 . A universal latent embedding space for sentences is first pre-trained on large text corpus, …
-
DEBERTA: DECODING-ENHANCED BERT WITH DISENTANGLED ATTENTION
2021 · International Conference on Learning Representations
Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks. In this paper we propose a new model architecture \textbf{DeBERTa} (\textbf{D}ecoding-\textbf{e}nhanced \textbf{BERT} with disentangled \textbf{a}ttention) that …
-
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
2021 · arXiv (Cornell University)
This paper presents a new pre-trained language model, DeBERTaV3, which improves the original DeBERTa model by replacing mask language modeling (MLM) with replaced token detection (RTD), a more sample-efficient pre-training task. Our analysis shows that …