SOTAVerified

DIRECT: Direct and Indirect Responses in Conversational Text Corpus

2021-11-01Findings (EMNLP) 2021Code Available1· sign in to hype

Junya Takayama, Tomoyuki Kajiwara, Yuki Arase

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

We create a large-scale dialogue corpus that provides pragmatic paraphrases to advance technology for understanding the underlying intentions of users. While neural conversation models acquire the ability to generate fluent responses through training on a dialogue corpus, previous corpora have mainly focused on the literal meanings of utterances. However, in reality, people do not always present their intentions directly. For example, if a person said to the operator of a reservation service “I don’t have enough budget.”, they, in fact, mean “please find a cheaper option for me.” Our corpus provides a total of 71,498 indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset. In addition, we propose three tasks to benchmark the ability of models to recognize and generate indirect and direct utterances. We also investigated the performance of state-of-the-art pre-trained models as baselines.

Reproductions