Question Answering

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 3576–3600 of 10817 papers

Title	Date	Tasks	Status
Double Retrieval and Ranking for Accurate Question Answering	Jan 16, 2022	Answer SelectionQuestion Answering	—Unverified
Expediting and Elevating Large Language Model Reasoning via Hidden Chain-of-Thought Decoding	Sep 13, 2024	Contrastive LearningLanguage Modeling	—Unverified
Categorizing Concepts With Basic Level for Vision-to-Language	Jun 1, 2018	ClusteringImage Captioning	—Unverified
Do Transformers Dream of Inference, or Can Pretrained Generative Models Learn Implicit Inferential Rules?	Nov 1, 2020	Multi-hop Question AnsweringQuestion Answering	—Unverified
Experiments on Hybrid Corpus-Based Sentiment Lexicon Acquisition	Apr 1, 2012	Document ClassificationQuestion Answering	—Unverified
Experiments with Easy-first nonprojective constituent parsing	Aug 1, 2014	Dependency ParsingMachine Translation	—Unverified
Expert Finding in Community Question Answering: A Review	Apr 21, 2018	Community Question AnsweringEnsemble Learning	—Unverified
Annotation Methodologies for Vision and Language Dataset Creation	Jul 10, 2016	Action RecognitionImage Description	—Unverified
Do Transformer Networks Improve the Discovery of Rules from Text?	Jun 1, 2022	Language ModelingLanguage Modelling	—Unverified
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering	May 17, 2025	Document RankingLarge Language Model	—Unverified
A Framework for the Classification and Annotation of Multiword Expressions in Dialectal Arabic	Oct 1, 2014	Entity Extraction using GANGeneral Classification	—Unverified
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey	Dec 3, 2024	Cross-Modal RetrievalNatural Language Understanding	—Unverified
Explainable Artificial Intelligence Recommendation System by Leveraging the Semantics of Adverse Childhood Experiences: Proof-of-Concept Prototype Development	Nov 6, 2020	Explainable artificial intelligenceGraph Generation	—Unverified
Font-Agent: Enhancing Font Understanding with Large Language Models	Jan 1, 2025	Font GenerationQuestion Answering	—Unverified
Explainable Assessment of Healthcare Articles with QA	May 1, 2022	ArticlesExplanation Generation	—Unverified
DoT: An efficient Double Transformer for NLP tasks with tables	Jun 1, 2021	Question Answering	—Unverified
Do Smaller Language Models Answer Contextualised Questions Through Memorisation Or Generalisation?	Nov 21, 2023	Question AnsweringSemantic Similarity	—Unverified
Explainable Fact-checking through Question Answering	Oct 11, 2021	Decision MakingFact Checking	—Unverified
A RAG-based Question Answering System Proposal for Understanding Islam: MufassirQAS LLM	Jan 27, 2024	ArticlesChatbot	—Unverified
Annotation and Analysis of Discourse Relations, Temporal Relations and Multi-Layered Situational Relations in Japanese Texts	Dec 1, 2016	ArticlesNatural Language Inference	—Unverified
Case-Based Abductive Natural Language Inference	Sep 30, 2020	Natural Language InferenceQuestion Answering	—Unverified
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?	Apr 5, 2024	Question AnsweringRecommendation Systems	—Unverified
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures	Feb 23, 2024	Question AnsweringText Generation	—Unverified
A Concept-Centric Approach to Multi-Modality Learning	Dec 18, 2024	Image-text matchingQuestion Answering	—Unverified
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment	Jul 1, 2023	Language ModelingLanguage Modelling	—Unverified

Show:10 25 50

← PrevPage 144 of 433Next →

All datasets SQuAD2.0 SQuAD1.1 HotpotQA PIQA BoolQ COPA TriviaQA SQuAD1.1 dev Natural Questions OpenBookQA TruthfulQA MultiRC

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	IE-Net (ensemble)	EM	90.94	—	Unverified
2	FPNet (ensemble)	EM	90.87	—	Unverified
3	IE-NetV2 (ensemble)	EM	90.86	—	Unverified
4	SA-Net on Albert (ensemble)	EM	90.72	—	Unverified
5	SA-Net-V2 (ensemble)	EM	90.68	—	Unverified
6	FPNet (ensemble)	EM	90.6	—	Unverified
7	Retro-Reader (ensemble)	EM	90.58	—	Unverified
8	EntitySpanFocusV2 (ensemble)	EM	90.52	—	Unverified
9	TransNets + SFVerifier + SFEnsembler (ensemble)	EM	90.49	—	Unverified
10	EntitySpanFocus+AT (ensemble)	EM	90.45	—	Unverified