Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 3851–3860 of 10817 papers

Title	Date	Tasks	Status
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms	Nov 17, 2024	DiagnosticMiscellaneous	—Unverified
Domain Adaptation with Active Learning for Coreference Resolution	Apr 1, 2014	Active Learningcoreference-resolution	—Unverified
Domain Adaptation of VLM for Soccer Video Understanding	May 20, 2025	Action ClassificationDomain Adaptation	—Unverified
Few-shot Question Generation for Personalized Feedback in Intelligent Tutoring Systems	Jun 8, 2022	Generative Question AnsweringQuestion Answering	—Unverified
Beyond Attention: Toward Machines with Intrinsic Higher Mental States	May 2, 2025	Question Answering	—Unverified
Annotating Educational Questions for Student Response Analysis	May 1, 2018	Question AnsweringWord Embeddings	—Unverified
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches	Mar 17, 2024	Image CaptioningQuestion Answering	—Unverified
FFA Sora, video generation as fundus fluorescein angiography simulator	Dec 23, 2024	Privacy PreservingQuestion Answering	—Unverified
Dolphin: A Challenging and Diverse Benchmark for Arabic NLG	May 24, 2023	Dialogue GenerationDiversity	—Unverified
A Flexible, Efficient and Accurate Framework for Community Question Answering Pipelines	Jul 1, 2018	Community Question AnsweringQuestion Answering	—Unverified

Show:10 25 50

← PrevPage 386 of 1082Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified