Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 4526–4550 of 10817 papers

Title	Date	Tasks	Status
Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT	Apr 11, 2023	DiagnosticImage Captioning	—Unverified
A Comparative Evaluation of Visual and Natural Language Question Answering Over Linked Data	Jul 19, 2019	Natural Language QueriesQuestion Answering	—Unverified
Information Extraction over Structured Data: Question Answering with Freebase	Jun 1, 2014	Information RetrievalQuestion Answering	—Unverified
Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains	Oct 24, 2024	Knowledge GraphsQuestion Answering	—Unverified
Automatic Rule Extraction from Long Short Term Memory Networks	Feb 8, 2017	Question AnsweringSentiment Analysis	—Unverified
Information Discovery in e-Commerce	Oct 8, 2024	Information RetrievalKnowledge Graphs	—Unverified
An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts	Jan 21, 2025	Question AnsweringSentiment Analysis	—Unverified
Automatic recognition of habituals: a three-way classification of clausal aspect	Sep 1, 2015	General ClassificationQuestion Answering	—Unverified
Advancing Large Language Model Attribution through Self-Improving	Oct 17, 2024	Language ModelingLanguage Modelling	—Unverified
Information Extraction From Co-Occurring Similar Entities	Feb 10, 2021	DescriptiveKnowledge Graphs	—Unverified
Automatic Question Generation using Relative Pronouns and Adverbs	Jul 1, 2018	DescriptiveDialogue Generation	—Unverified
Decision Knowledge Graphs: Construction of and Usage in Question Answering for Clinical Practice Guidelines	Aug 6, 2023	Knowledge GraphsQuestion Answering	—Unverified
Decipherment	Aug 1, 2013	DeciphermentPart-Of-Speech Tagging	—Unverified
Automatic Question-Answering Using A Deep Similarity Neural Network	Aug 5, 2017	Question Answering	—Unverified
An Empirical Evaluation of Visual Question Answering for Novel Objects	Apr 8, 2017	Question AnsweringVisual Question Answering	—Unverified
Information Extraction from Documents: Question Answering vs Token Classification in real-world setups	Apr 21, 2023	ClassificationFew-Shot Learning	—Unverified
Information Gathering in Networks via Active Exploration	Apr 24, 2015	Experimental DesignInformativeness	—Unverified
Informing clinical assessment by contextualizing post-hoc explanations of risk prediction models in type-2 diabetes	Feb 11, 2023	Question Answering	—Unverified
Deceptive Answer Prediction with User Preference Graph	Aug 1, 2013	Answer SelectionCommunity Question Answering	—Unverified
Deception Detection in News Reports in the Russian Language: Lexics and Discourse	Sep 1, 2017	Deception DetectionFact Checking	—Unverified
Automatic Question Answering for Medical MCQs: Can It go Further than Information Retrieval?	Sep 1, 2019	Information RetrievalMultiple-choice	—Unverified
Automatic Question-Answer Generation for Long-Tail Knowledge	Mar 3, 2024	Answer GenerationKnowledge Graphs	—Unverified
An Empirical Evaluation of various Deep Learning Architectures for Bi-Sequence Classification Tasks	Jul 17, 2016	ClassificationDeep Learning	—Unverified
An Empirical Evaluation of Large Language Models on Consumer Health Questions	Dec 31, 2024	Medical Question AnsweringQuestion Answering	—Unverified
DeCAP: Context-Adaptive Prompt Generation for Debiasing Zero-shot Question Answering in Large Language Models	Mar 25, 2025	FairnessQuestion Answering	—Unverified

Show:10 25 50

← PrevPage 182 of 433Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified