SOTAVerified

CSD: A Chinese Dataset for Subtext Problem

2022-01-16ACL ARR January 2022Code Available0· sign in to hype

Anonymous

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

Subtext is a kind of deep semantics which can be acquired after one or more rounds of expression transformation. As a popular way of expressing one's intentions, it is well worth studying. In this paper, we propose two subtext-related tasks which are termed ``subtext recognition'' and ``subtext recovery'' and make a clear definition for their purposes. Moreover, we build a Chinese dataset whose source data comes from popular social media (e.g. Weibo, Netease Music, Zhihu, and Bilibili) and propose a new evaluation metric termed ``Two-stages Annotation Evaluation'' (TAE) for the validation of a multi-turn annotation process.

Reproductions