Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 676–700 of 10817 papers

Title	Date	Tasks	Status	Hype
InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models	Jan 19, 2025	BenchmarkingQuestion Answering	CodeCode Available	1
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning	Jan 13, 2025	Causal DiscoveryCausal Inference	CodeCode Available	1
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring	Jan 9, 2025	Question Answering	CodeCode Available	1
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark	Jan 9, 2025	FairnessHallucination	CodeCode Available	1
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models	Jan 9, 2025	BenchmarkingMathematical Problem-Solving	CodeCode Available	1
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation	Jan 6, 2025	Language Model EvaluationLanguage Modeling	CodeCode Available	1
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?	Jan 5, 2025	Image CaptioningImage to text	CodeCode Available	1
Predicting the Performance of Black-box LLMs through Self-Queries	Jan 2, 2025	Question Answering	CodeCode Available	1
Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question Answering	Jan 1, 2025	Large Language ModelMultimodal Large Language Model	CodeCode Available	1
Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner	Dec 30, 2024	Question AnsweringTable Recognition	CodeCode Available	1
Long Context vs. RAG for LLMs: An Evaluation and Revisits	Dec 27, 2024	Question AnsweringRAG	CodeCode Available	1
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions	Dec 27, 2024	Human-Object Interaction DetectionObject	CodeCode Available	1
Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation	Dec 24, 2024	Graph Question AnsweringHallucination	CodeCode Available	1
CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era	Dec 24, 2024	Knowledge Base Question AnsweringKnowledge Graphs	CodeCode Available	1
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating	Dec 24, 2024	document understandingQuestion Answering	CodeCode Available	1
Property Enhanced Instruction Tuning for Multi-task Molecule Generation with Large Language Models	Dec 24, 2024	Machine TranslationMolecular Property Prediction	CodeCode Available	1
Resource-Aware Arabic LLM Creation: Model Adaptation, Integration, and Multi-Domain Testing	Dec 23, 2024	ArabicMMLUDialect Identification	CodeCode Available	1
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding	Dec 21, 2024	AttributeQuestion Answering	CodeCode Available	1
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization	Dec 19, 2024	Contrastive LearningDecision Making	CodeCode Available	1
Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question Answering	Dec 18, 2024	graph constructionknowledge editing	CodeCode Available	1
MedCoT: Medical Chain of Thought via Hierarchical Expert	Dec 18, 2024	DiagnosticMedical Visual Question Answering	CodeCode Available	1
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation	Dec 17, 2024	Question AnsweringRAG	CodeCode Available	1
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants	Dec 17, 2024	Image CaptioningQuestion Answering	CodeCode Available	1
SCITAT: A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types	Dec 16, 2024	Question Answering	CodeCode Available	1
UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models	Dec 16, 2024	Question Answering	CodeCode Available	1

Show:10 25 50

← PrevPage 28 of 433Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified