SOTAVerified

Reading Comprehension

Most current question answering datasets frame the task as reading comprehension where the question is about a paragraph or document and the answer often is a span in the document.

Some specific tasks of reading comprehension include multi-modal machine reading comprehension and textual machine reading comprehension, among others. In the literature, machine reading comprehension can be divide into four categories: cloze style, multiple choice, span prediction, and free-form answer. Read more about each category here.

Benchmark datasets used for testing a model's reading comprehension abilities include MovieQA, ReCoRD, and RACE, among others.

The Machine Reading group at UCL also provides an overview of reading comprehension tasks.

Figure source: A Survey on Machine Reading Comprehension: Tasks, Evaluation Metrics and Benchmark Datasets

Papers

Showing 351–400 of 1760 papers

TitleStatusHype
Advances in Multi-turn Dialogue Comprehension: A Survey—0
Composing RNNs and FSTs for Small Data: Recovering Missing Characters in Old Hawaiian Text—0
Composing RNNs and FSTs for Small Data: Recovering Missing Characters in Old Hawaiian Text—0
A Survey on Machine Reading Comprehension: Tasks, Evaluation Metrics and Benchmark Datasets—0
A Comparative Study of Word Embeddings for Reading Comprehension—0
Deriving Commonsense Inference Tasks from Interactive Fictions—0
Detecting Causes of Stock Price Rise and Decline by Machine Reading Comprehension with BERT—0
DIFM:An effective deep interaction and fusion model for sentence matching—0
A Survey on Machine Reading Comprehension Systems—0
A Survey on Explainability in Machine Reading Comprehension—0
Advancements and Challenges in Bangla Question Answering Models: A Comprehensive Review—0
A Study on Contextualized Language Modeling for Machine Reading Comprehension—0
Comparative Analysis of Neural QA models on SQuAD—0
Commonsense Knowledge + BERT for Level 2 Reading Comprehension Ability Test—0
2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval—0
A study of Vietnamese readability assessing through semantic and statistical features—0
Commonsense Inference in Natural Language Processing (COIN) - Shared Task Report—0
A Study of the Tasks and Models in Machine Reading Comprehension—0
Analyzing Multiple-Choice Reading and Listening Comprehension Tests—0
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction—0
Commonsense Evidence Generation and Injection in Reading Comprehension—0
A Strong Lexical Matching Method for the Machine Comprehension Test—0
Commonsense knowledge adversarial dataset that challenges ELECTRA—0
Commonsense Knowledge Base Completion and Generation—0
Analyzing and Mitigating Interference in Neural Architecture Search—0
CoMiC: Adapting a Short Answer Assessment System for Answer Selection—0
A Survey of Machine Narrative Reading Comprehension Assessments—0
Complementary Advantages of ChatGPTs and Human Readers in Reasoning: Evidence from English Text Reading Comprehension—0
Complex Factoid Question Answering with a Free-Text Knowledge Graph—0
Complex Reading Comprehension Through Question Decomposition—0
Complex Word Identification Based on Frequency in a Learner Corpus—0
Composing Answer from Multi-spans for Reading Comprehension—0
CoMeT: Integrating different levels of linguistic modeling for meaning assessment—0
Assessing the Benchmarking Capacity of Machine Reading Comprehension Datasets—0
A Survey on Measuring and Mitigating Reasoning Shortcuts in Machine Reading Comprehension—0
Comprehending Knowledge Graphs with Large Language Models for Recommender Systems—0
Comprehensive Multi-Dataset Evaluation of Reading Comprehension—0
Compressing Long Context for Enhancing RAG with AMR-based Concept Distillation—0
A Discriminative Model for Identifying Readers and Assessing Text Comprehension from Eye Movements—0
Computational Approaches to Sentence Completion—0
Computing Semantic Text Similarity Using Rich Features—0
Deep Understanding based Multi-Document Machine Reading Comprehension—0
Analyzing Zero-shot Cross-lingual Transfer in Supervised NLP Tasks—0
DEIM: An effective deep encoding and interaction model for sentence matching—0
Combining Probabilistic Logic and Deep Learning for Self-Supervised Learning—0
Combining Formal and Distributional Models of Temporal and Intensional Semantics—0
Constructing Datasets for Multi-hop Reading Comprehension Across Documents—0
Assessing Distractors in Multiple-Choice Tests—0
Collecting high-quality adversarial data for machine reading comprehension tasks with humans and models in the loop—0
Assessing Conformance of Manually Simplified Corpora with User Requirements: the Case of Autistic Readers—0
Show:102550
← PrevPage 8 of 36Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Rational Reasoner / IDOLTest80.6—Unverified
2AMR-LE-EnsembleTest80—Unverified
3MERIt(MERIt-deberta-v2-xxlarge )Test79.3—Unverified
4MERIt-deberta-v2-xxlarge deberta.v2.xxlarge.path.override_True.norm_1.1.0.w2.A100.cp200.s42Test79.3—Unverified
5Knowledge modelTest79.2—Unverified
6DeBERTa-v2-xxlarge-AMR-LE-ContrapositionTest77.2—Unverified
7LReasoner ensembleTest76.1—Unverified
8ELECTRA and ALBERTTest71—Unverified
9WWZTest69.7—Unverified
10xlnet-large-uncased [extended data]Test69.3—Unverified
#ModelMetricClaimedVerifiedStatus
1ALBERT (Ensemble)Accuracy91.4—Unverified
2Megatron-BERT (ensemble)Accuracy90.9—Unverified
3ALBERTxxlarge+DUMA(ensemble)Accuracy89.8—Unverified
4Megatron-BERTAccuracy89.5—Unverified
5XLNetAccuracy (Middle)88.6—Unverified
6DeBERTalargeAccuracy86.8—Unverified
7B10-10-10Accuracy85.7—Unverified
8RoBERTaAccuracy83.2—Unverified
9Orca 2-13BAccuracy82.87—Unverified
10Orca 2-7BAccuracy80.79—Unverified
#ModelMetricClaimedVerifiedStatus
1Golden TransformerAverage F10.94—Unverified
2MT5 LargeAverage F10.84—Unverified
3ruRoberta-large finetuneAverage F10.83—Unverified
4ruT5-large-finetuneAverage F10.82—Unverified
5Human BenchmarkAverage F10.81—Unverified
6ruT5-base-finetuneAverage F10.77—Unverified
7ruBert-large finetuneAverage F10.76—Unverified
8ruBert-base finetuneAverage F10.74—Unverified
9RuGPT3XL few-shotAverage F10.74—Unverified
10RuGPT3LargeAverage F10.73—Unverified
#ModelMetricClaimedVerifiedStatus
1RoBERTa-LargeOverall: F164.4—Unverified
2BERT-LargeOverall: F162.7—Unverified
3BiDAFOverall: F128.5—Unverified
#ModelMetricClaimedVerifiedStatus
1BERTMSE0.05—Unverified
#ModelMetricClaimedVerifiedStatus
1BERT pretrained on MIMIC-IIIAnswer F163.55—Unverified