Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 4051–4075 of 10817 papers

Title	Date	Tasks	Status
From Parse-Execute to Parse-Execute-Refine: Improving Semantic Parser for Complex Question Answering over Knowledge Base	May 5, 2023	Knowledge Base Question AnsweringQuestion Answering	—Unverified
Coal Mining Question Answering with LLMs	Oct 3, 2024	Prompt EngineeringQuestion Answering	—Unverified
From Pixels to Objects: Cubic Visual Attention for Visual Question Answering	Jun 4, 2022	ObjectQuestion Answering	—Unverified
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing	Nov 5, 2024	Change DetectionContrastive Learning	—Unverified
Hard to Cheat: A Turing Test based on Answering Questions about Images	Jan 14, 2015	Question Answering	—Unverified
Do Explanations make VQA Models more Predictable to a Human?	Oct 29, 2018	Question AnsweringVisual Question Answering	—Unverified
From Questions to Insightful Answers: Building an Informed Chatbot for University Resources	May 13, 2024	ChatbotLanguage Modeling	—Unverified
From RAGs to rich parameters: Probing how language models utilize external knowledge over parametric information for factual queries	Jun 18, 2024	Question AnsweringRAG	—Unverified
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs	Jun 5, 2025	cross-modal alignmentDense Captioning	—Unverified
Better Query Graph Selection for Knowledge Base Question Answering	Apr 27, 2022	Knowledge Base Question AnsweringQuestion Answering	—Unverified
Affordances in Grounded Language Learning	Jul 1, 2018	Grounded language learningQuestion Answering	—Unverified
From Retrieval to Generation: Comparing Different Approaches	Feb 27, 2025	Language ModelingLanguage Modelling	—Unverified
A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry	Apr 24, 2024	Information RetrievalLanguage Modeling	—Unverified
From Shakespeare to Twitter: What are Language Styles all about?	Sep 1, 2017	AllLexical Simplification	—Unverified
From Shallow to Deep: Compositional Reasoning over Graphs for Visual Question Answering	Jun 25, 2022	Question AnsweringVisual Question Answering	—Unverified
From Strings to Things: Knowledge-Enabled VQA Model That Can Read and Reason	Oct 1, 2019	Graph Neural NetworkQuestion Answering	—Unverified
Harnessing AI for efficient analysis of complex policy documents: a case study of Executive Order 14110	Jun 10, 2024	Question Answering	—Unverified
From text to multimodal: a survey of adversarial example generation in question answering systems	Dec 26, 2023	Question AnsweringQuestion Generation	—Unverified
From Text to Visuals: Using LLMs to Generate Math Diagrams with Vector Graphics	Mar 10, 2025	MathQuestion Answering	—Unverified
From Textual Entailment to Knowledgeable Machines	Nov 1, 2013	Natural Language InferenceQuestion Answering	—Unverified
HARPY: Hypernyms and Alignment of Relational Paraphrases	Aug 1, 2014	Question Answering	—Unverified
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance	Jun 26, 2020	Decision MakingQuestion Answering	—Unverified
Does the "most sinfully decadent cake ever" taste good? Answering Yes/No Questions from Figurative Contexts	Sep 24, 2023	Question Answering	—Unverified
From Visual to Acoustic Question Answering	Feb 28, 2019	Acoustic Question AnsweringPosition	—Unverified
Better Early than Late: Fusing Topics with Word Embeddings for Neural Question Paraphrase Identification	Jul 22, 2020	Community Question AnsweringParaphrase Identification	—Unverified

Show:10 25 50

← PrevPage 163 of 433Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified