DebateSum: A large-scale argument mining and summarization dataset

2020-11-14COLING (ArgMining) 2020Code Available1· sign in to hype

Allen Roush, Arvind Balaji

Code Available — Be the first to reproduce this paper.

Code

github.com/Hellisotherpeople/DebateSum
OfficialIn papernone★ 55
github.com/arvind-balaji/debate-cards
Officialnone★ 40
github.com/Hellisotherpeople/debate2vec
OfficialIn paperpytorch★ 4

Abstract

Prior work in Argument Mining frequently alludes to its potential applications in automatic debating systems. Despite this focus, almost no datasets or models exist which apply natural language processing techniques to problems found within competitive formal debate. To remedy this, we present the DebateSum dataset. DebateSum consists of 187,386 unique pieces of evidence with corresponding argument and extractive summaries. DebateSum was made using data compiled by competitors within the National Speech and Debate Association over a 7-year period. We train several transformer summarization models to benchmark summarization performance on DebateSum. We also introduce a set of fasttext word-vectors trained on DebateSum called debate2vec. Finally, we present a search engine for this dataset which is utilized extensively by members of the National Speech and Debate Association today. The DebateSum search engine is available to the public here: http://www.debate.cards

Tasks

Abstractive Text Summarization Argument Mining Document Summarization Extractive Text Summarization Information Retrieval Query-Based Extractive Summarization Text Summarization

Benchmark Results

Dataset	Model	Metric	Claimed	Verified	Status
DebateSum	Longformer-Base	ROUGE-L	57.21	—	Unverified
DebateSum	GPT2-Medium	ROUGE-L	53.23	—	Unverified
DebateSum	BERT-Large	ROUGE-L	49.98	—	Unverified

DebateSum: A large-scale argument mining and summarization dataset

Code

Abstract

Tasks

Benchmark Results

Reproductions