Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 4541–4550 of 10817 papers

Title	Date	Tasks	Status
Information Extraction from Documents: Question Answering vs Token Classification in real-world setups	Apr 21, 2023	ClassificationFew-Shot Learning	—Unverified
Information Gathering in Networks via Active Exploration	Apr 24, 2015	Experimental DesignInformativeness	—Unverified
Informing clinical assessment by contextualizing post-hoc explanations of risk prediction models in type-2 diabetes	Feb 11, 2023	Question Answering	—Unverified
Deceptive Answer Prediction with User Preference Graph	Aug 1, 2013	Answer SelectionCommunity Question Answering	—Unverified
Deception Detection in News Reports in the Russian Language: Lexics and Discourse	Sep 1, 2017	Deception DetectionFact Checking	—Unverified
Automatic Question Answering for Medical MCQs: Can It go Further than Information Retrieval?	Sep 1, 2019	Information RetrievalMultiple-choice	—Unverified
Automatic Question-Answer Generation for Long-Tail Knowledge	Mar 3, 2024	Answer GenerationKnowledge Graphs	—Unverified
An Empirical Evaluation of various Deep Learning Architectures for Bi-Sequence Classification Tasks	Jul 17, 2016	ClassificationDeep Learning	—Unverified
An Empirical Evaluation of Large Language Models on Consumer Health Questions	Dec 31, 2024	Medical Question AnsweringQuestion Answering	—Unverified
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models	Mar 25, 2025	FairnessQuestion Answering	—Unverified

Show:10 25 50

← PrevPage 455 of 1082Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified