SOTAVerified

Open-Domain Question Answering

Open-domain question answering is the task of question answering on open-domain datasets such as Wikipedia.

Papers

Showing 150 of 494 papers

TitleStatusHype
TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document ReasoningCode2
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k0
ECoRAG: Evidentiality-guided Compression for Long Context RAGCode1
GenKI: Enhancing Open-Domain Question Answering with Knowledge Integration and Controllable Generation in Large Language ModelsCode0
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement LearningCode1
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization0
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation0
Scaling Reasoning can Improve Factuality in Large Language ModelsCode0
Benchmarking LLM-based Relevance Judgment MethodsCode0
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task0
CoRAG: Collaborative Retrieval-Augmented Generation0
Dense Passage Retrieval in Conversational SearchCode0
Knowledge-Aware Iterative Retrieval for Multi-Agent Systems0
Optimizing open-domain question answering with graph-based retrieval augmented generation0
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering0
WebFAQ: A Multilingual Collection of Natural Q&A Datasets for Dense Retrieval0
From Retrieval to Generation: Comparing Different Approaches0
Few-Shot Multilingual Open-Domain QA from 5 ExamplesCode0
RA-MTR: A Retrieval Augmented Multi-Task Reader based Approach for Inspirational Quote Extraction from Long DocumentsCode0
RoseRAG: Robust Retrieval-augmented Generation with Small-scale LLMs via Margin-aware Preference Optimization0
Talk Structurally, Act Hierarchically: A Collaborative Framework for LLM Multi-Agent SystemsCode2
ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval0
Passage Segmentation of Documents for Extractive Question Answering0
Parallel Key-Value Cache Fusion for Position Invariant RAG0
WebWalker: Benchmarking LLMs in Web TraversalCode11
Improving Generated and Retrieved Knowledge Combination Through Zero-shot Generation0
Accelerating Manufacturing Scale-Up from Material Discovery Using Agentic Web Navigation and Retrieval-Augmented AI for Process Engineering Schematics Design0
DynRank: Improving Passage Retrieval with Dynamic Zero-Shot Prompting Based on Question Classification0
Context Awareness Gate For Retrieval Augmented GenerationCode1
Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering0
Invar-RAG: Invariant LLM-aligned Retrieval for Better Generation0
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking0
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation EngineeringCode2
Improve Dense Passage Retrieval with Entailment Tuning0
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via CompressionCode1
Advancing Large Language Model Attribution through Self-Improving0
Open Domain Question Answering with Conflicting Contexts0
BanglaQuAD: A Bengali Open-domain Question Answering Dataset0
LoRE: Logit-Ranked Retriever Ensemble for Enhancing Open-Domain Question Answering0
Retriever-and-Memory: Towards Adaptive Note-Enhanced Retrieval-Augmented GenerationCode2
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation SystemsCode0
Detecting Temporal Ambiguity in QuestionsCode0
Exploring Hint Generation Approaches in Open-Domain Question AnsweringCode1
A Multimodal Dense Retrieval Approach for Speech-Based Open-Domain Question Answering0
A Statistical Framework for Data-dependent Retrieval-Augmented Models0
Towards Human-Level Understanding of Complex Process Engineering Schematics: A Pedagogical, Introspective Multi-Agent Framework for Open-Domain Question Answering0
W-RAG: Weakly Supervised Dense Retrieval in RAG for Open-domain Question AnsweringCode1
FastFiD: Improve Inference Efficiency of Open Domain Question Answering via Sentence SelectionCode1
Enhancing Robustness of Retrieval-Augmented Language Models with In-Context Learning0
Adaptive Contrastive Decoding in Retrieval-Augmented Generation for Handling Noisy Contexts0
Show:102550
← PrevPage 1 of 10Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1somebodyKILT-RL2.62Unverified
2WikipediaKILT-RL2.46Unverified
3arxiv.org/abs/2103.06332KILT-RL2.36Unverified
4BART + DPRKILT-RL1.9Unverified
5RAGKILT-RL1.69Unverified
6T5-baseKILT-RL0Unverified
7GENREKILT-RL0Unverified
8Multi-task DPRKILT-RL0Unverified
9BARTKILT-RL0Unverified
10Training Set Retrieval (top 1)KILT-RL0Unverified
#ModelMetricClaimedVerifiedStatus
1Re2GKILT-EM43.56Unverified
2intersectKILT-EM38.78Unverified
3KGI_0KILT-EM36.36Unverified
4WikipediaKILT-EM35.32Unverified
5RAGKILT-EM32.69Unverified
6BERT + DPRKILT-EM31.99Unverified
7BART + DPRKILT-EM30.06Unverified
8Multitask DPR + BARTKILT-EM29.09Unverified
9SphereKILT-EM0Unverified
10T5-baseKILT-EM0Unverified
#ModelMetricClaimedVerifiedStatus
1Re2GKILT-EM57.91Unverified
2intersectKILT-EM50.56Unverified
3WikipediaKILT-EM45.55Unverified
4KGI_0KILT-EM42.85Unverified
5Multitask DPR + BARTKILT-EM42.36Unverified
6RAGKILT-EM38.13Unverified
7BERT + DPRKILT-EM34.48Unverified
8BART + DPRKILT-EM31.4Unverified
9Multi-task DPRKILT-EM0Unverified
10SphereKILT-EM0Unverified
#ModelMetricClaimedVerifiedStatus
1intersectKILT-EM18.06Unverified
2WikipediaKILT-EM11.71Unverified
3Multitask DPR + BARTKILT-EM9.53Unverified
4RAGKILT-EM3.21Unverified
5BART + DPRKILT-EM1.96Unverified
6BERT + DPRKILT-EM0.74Unverified
7SphereKILT-EM0Unverified
8Multi-task DPRKILT-EM0Unverified
9GENREKILT-EM0Unverified
10chriskueiKILT-EM0Unverified
#ModelMetricClaimedVerifiedStatus
1SpanBERTF184.8Unverified
2Cluster-Former (#C=512)EM68Unverified
3Locality-Sensitive HashingEM66Unverified
4Multi-passage BERTEM65.1Unverified
5Sparse AttentionEM64.7Unverified
6DECAPROPEM62.2Unverified
7Bi-Attention + DCU-LSTMN-gram F159.5Unverified
8Denoising QAEM58.8Unverified
9DecaPropEM56.8Unverified
10AMANDAN-gram F156.6Unverified
#ModelMetricClaimedVerifiedStatus
1Fourier TransformerRouge-L26.9Unverified
2QGRouge-L26.4Unverified
3BARTRouge-L24.3Unverified
4E-MCARouge-L24Unverified
5Transformer Multitask + LayerDropRouge-L23.4Unverified
6Multi-InrerleaveRouge-L14.63Unverified
#ModelMetricClaimedVerifiedStatus
1Evidence Aggregation via R^3 Re-RankingEM (Quasar-T)42.3Unverified
2Denoising QAEM (Quasar-T)42.2Unverified
3DecaPropEM (Quasar-T)38.6Unverified
4R^3EM (Quasar-T)35.3Unverified
5GAEM (Quasar-T)26.4Unverified
6BiDAFEM (Quasar-T)25.9Unverified
#ModelMetricClaimedVerifiedStatus
1FiEExact Match58.4Unverified
2R2-D2 HN-DPRExact Match55.9Unverified
3UniK-QAExact Match54.9Unverified
4UnitedQA (Hybrid)Exact Match54.7Unverified
5BPR (linear scan; l=1000)Exact Match41.6Unverified
#ModelMetricClaimedVerifiedStatus
1SPARTAEM59.3Unverified
2Blended RAGEM57.63Unverified
3BERTseriniEM50.2Unverified
4BERTseriniEM38.6Unverified
#ModelMetricClaimedVerifiedStatus
1UniK-QAExact Match57.7Unverified
2FiE+PAQExact Match56.3Unverified
3FiEExact Match52.4Unverified
4EMDR2Exact Match48.7Unverified
#ModelMetricClaimedVerifiedStatus
1DrQAEM70Unverified
2DCNEM66.2Unverified
3MPCMEM65.5Unverified
#ModelMetricClaimedVerifiedStatus
1ERNIE 2.0 LargeEM64.2Unverified
2ERNIE 2.0 BaseEM61.3Unverified
#ModelMetricClaimedVerifiedStatus
1UniK-QAExact Match65.5Unverified
2BPR (linear scan; l=1000)Exact Match56.8Unverified
#ModelMetricClaimedVerifiedStatus
1EMDR2Exact Match52.5Unverified
#ModelMetricClaimedVerifiedStatus
1UnitedQA (Hybrid)Exact Match70.5Unverified