conference-paper

Iterative policy learning in end-to-end trainable task-oriented neural dialog models

  • 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Research footprint

At a glance

الاستشهادات
94
المراجع
55
Comments
0
Paper overview

Abstract

In this paper, we present a deep reinforcement learning (RL) framework for iterative dialog policy optimization in end-to-end task-oriented dialog systems. Popular approaches in learning dialog policy with RL include letting a dialog agent to learn against a user simulator. Building a reliable user simulator, however, is not trivial, often as difficult as building a good dialog agent. We address this challenge by jointly optimizing the dialog agent and the user simulator with deep RL by simulating dialogs between the two agents. We first bootstrap a basic dialog agent and a basic user simulator by learning directly from dialog corpora with supervised training. We then improve them further by letting the two agents to conduct task-oriented dialogs and iteratively optimizing their policies with deep RL. Both the dialog agent and the user simulator are designed with neural network models that can be trained end-to-end. Our experiment results show that the proposed method leads to promising improvements on task success rate and total task reward comparing to supervised training and single-agent RL training baseline models.

Record transparency

Publication details

DOI
10.1109/asru.2017.8268975
OpenAlex
W2963043030
Document type
conference-paper
Language
EN
Source
2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.