conference-paper Open access

On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark

  • Findings of the Association for Computational Linguistics: ACL 2022
Research footprint

At a glance

Citations
49
References
79
Comments
0
Paper overview

Abstract

Dialogue safety problems severely limit the real-world deployment of neural conversational models and have attracted great research interests recently. However, dialogue safety problems remain under-defined and the corresponding dataset is scarce. We propose a taxonomy for dialogue safety specifically designed to capture unsafe behaviors in humanbot dialogue settings, with focuses on contextsensitive unsafety, which is under-explored in prior works. To spur research in this direction, we compile DIASAFETY, a dataset with rich context-sensitive unsafe examples. Experiments show that existing safety guarding tools fail severely on our dataset. As a remedy, we train a dialogue safety classifier to provide a strong baseline for context-sensitive dialogue unsafety detection. With our classifier, we perform safety evaluations on popular conversational models and show that existing dialogue systems still exhibit concerning contextsensitive safety problems.

Record transparency

Publication details

DOI
10.18653/v1/2022.findings-acl.308
OpenAlex
W3207604419
Document type
conference-paper
Language
EN
Source
Findings of the Association for Computational Linguistics: ACL 2022
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.