Cross-Lingual Speaker Identification from Weak Local Evidence

2022-01-16ACL ARR January 2022Unverified0· sign in to hype

Anonymous

Unverified — Be the first to reproduce this paper.

Abstract

Speaker identification, determining which character said each utterance in text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-based features to directly approach this task, but these approaches come with significant drawbacks, such as lack of contextual reasoning and poor cross-lingual generalization. In this work, we propose a speaker identification framework that addresses these issues. We first extract large-scale distant supervision signals in English via general-purpose tools and heuristics, and then apply these weakly-labeled instances with a focus on encouraging contextual reasoning to train a cross-lingual language model. We show that our final model outperforms the previous state-of-the-art methods on two English speaker identification benchmarks by 5.4\% in accuracy, as well as two Chinese speaker identification datasets by up to 4.7\%.

Tasks

Language Modeling Language Modelling Speaker Identification

Cross-Lingual Speaker Identification from Weak Local Evidence

Abstract

Tasks

Reproductions