Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 701–750 of 10817 papers

Title	Date	Tasks	Status	Hype
RETQA: A Large-Scale Open-Domain Tabular Question Answering Dataset for Real Estate Sector	Dec 13, 2024	In-Context LearningQuestion Answering	CodeCode Available	1
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors	Dec 12, 2024	Question Answering	CodeCode Available	1
IMPACT: A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents	Dec 10, 2024	Cross-Modal RetrievalImage Classification	CodeCode Available	1
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations	Dec 9, 2024	Language ModelingLanguage Modelling	CodeCode Available	1
RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts	Dec 7, 2024	Change DetectionImage Comprehension	CodeCode Available	1
KG-Retriever: Efficient Knowledge Indexing for Retrieval-Augmented Large Language Models	Dec 7, 2024	Multi-hop Question AnsweringNavigate	CodeCode Available	1
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds	Dec 7, 2024	Question Answering	CodeCode Available	1
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining	Dec 3, 2024	backdoor defenseComputational Efficiency	CodeCode Available	1
GraphOTTER: Evolving LLM-based Graph Reasoning for Complex Table Question Answering	Dec 2, 2024	Question Answering	CodeCode Available	1
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos	Dec 2, 2024	Question AnsweringVideo Understanding	CodeCode Available	1
PerLA: Perceptive 3D Language Assistant	Nov 29, 2024	Dense CaptioningGraph Neural Network	CodeCode Available	1
Cross-modal Information Flow in Multimodal Large Language Models	Nov 27, 2024	Question AnsweringVisual Question Answering	CodeCode Available	1
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format	Nov 27, 2024	Dense Video CaptioningGrounded Video Question Answering	CodeCode Available	1
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks	Nov 26, 2024	Contrastive LearningQuestion Answering	CodeCode Available	1
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning	Nov 25, 2024	HallucinationQuestion Answering	CodeCode Available	1
Context Awareness Gate For Retrieval Augmented Generation	Nov 25, 2024	Open-Domain Question AnsweringQuestion Answering	CodeCode Available	1
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai	Nov 23, 2024	DiversityQuestion Answering	CodeCode Available	1
Teaching VLMs to Localize Specific Objects from In-context Examples	Nov 20, 2024	ObjectObject Tracking	CodeCode Available	1
A Survey of Medical Vision-and-Language Applications and Their Techniques	Nov 19, 2024	Decision MakingDiagnostic	CodeCode Available	1
BackdoorMBTI: A Backdoor Learning Multimodal Benchmark Tool Kit for Backdoor Defense Evaluation	Nov 17, 2024	Action Recognitionbackdoor defense	CodeCode Available	1
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework	Nov 14, 2024	Question AnsweringRAG	CodeCode Available	1
Controllable Context Sensitivity and the Knob Behind It	Nov 11, 2024	Question AnsweringRetrieval-augmented Generation	CodeCode Available	1
DELIFT: Data Efficient Language model Instruction Fine Tuning	Nov 7, 2024	Language ModelingLanguage Modelling	CodeCode Available	1
MEG: Medical Knowledge-Augmented Large Language Models for Question Answering	Nov 6, 2024	Knowledge Graph EmbeddingsMultiple-choice	CodeCode Available	1
MILU: A Multi-task Indic Language Understanding Benchmark	Nov 4, 2024	Multiple-choiceQuestion Answering	CodeCode Available	1
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering	Nov 1, 2024	Medical Question AnsweringQuestion Answering	CodeCode Available	1
Birdie: Advancing State Space Models with Reward-Driven Objectives and Curricula	Nov 1, 2024	Computational EfficiencyQuestion Answering	CodeCode Available	1
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection	Oct 31, 2024	Change DetectionQuestion Answering	CodeCode Available	1
Nearest Neighbor Normalization Improves Multimodal Retrieval	Oct 31, 2024	Cross-Modal RetrievalImage Captioning	CodeCode Available	1
Distinguishing Ignorance from Error in LLM Hallucinations	Oct 29, 2024	HallucinationQuestion Answering	CodeCode Available	1
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning	Oct 25, 2024	AllComputational Efficiency	CodeCode Available	1
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations	Oct 24, 2024	Instruction FollowingQuestion Answering	CodeCode Available	1
Large Language Models Reflect the Ideology of their Creators	Oct 24, 2024	Question AnsweringText Summarization	CodeCode Available	1
Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective	Oct 23, 2024	graph constructionKnowledge Graphs	CodeCode Available	1
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning	Oct 23, 2024	Question AnsweringSpeech Recognition	CodeCode Available	1
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning	Oct 23, 2024	Image CaptioningInstruction Following	CodeCode Available	1
Progressive Compositionality In Text-to-Image Generative Models	Oct 22, 2024	AttributeContrastive Learning	CodeCode Available	1
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression	Oct 20, 2024	In-Context LearningLong-Context Understanding	CodeCode Available	1
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning	Oct 18, 2024	HallucinationKnowledge Base Question Answering	CodeCode Available	1
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems	Oct 18, 2024	BenchmarkingQuestion Answering	CodeCode Available	1
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines	Oct 16, 2024	Question AnsweringVisual Question Answering	CodeCode Available	1
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine	Oct 16, 2024	Language ModelingLanguage Modelling	CodeCode Available	1
RuleRAG: Rule-guided retrieval-augmented generation with language models for question answering	Oct 15, 2024	In-Context LearningInstruction Following	CodeCode Available	1
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models	Oct 14, 2024	2kBenchmarking	CodeCode Available	1
Towards Foundation Models for 3D Vision: How Close Are We?	Oct 14, 2024	Question AnsweringVisual Question Answering	CodeCode Available	1
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering	Oct 12, 2024	Answer GenerationBlocking	CodeCode Available	1
Skipping Computations in Multimodal LLMs	Oct 12, 2024	Question AnsweringVisual Question Answering	CodeCode Available	1
Dynamic Multimodal Evaluation with Flexible Complexity by Vision-Language Bootstrapping	Oct 11, 2024	MMEQuestion Answering	CodeCode Available	1
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models	Oct 11, 2024	Few-Shot LearningMultiple-choice	CodeCode Available	1
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models	Oct 10, 2024	Question AnsweringReinforcement Learning (RL)	CodeCode Available	1

Show:10 25 50

← PrevPage 15 of 217Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified