SOTAVerified

Reading Comprehension

Most current question answering datasets frame the task as reading comprehension where the question is about a paragraph or document and the answer often is a span in the document.

Some specific tasks of reading comprehension include multi-modal machine reading comprehension and textual machine reading comprehension, among others. In the literature, machine reading comprehension can be divide into four categories: cloze style, multiple choice, span prediction, and free-form answer. Read more about each category here.

Benchmark datasets used for testing a model's reading comprehension abilities include MovieQA, ReCoRD, and RACE, among others.

The Machine Reading group at UCL also provides an overview of reading comprehension tasks.

Figure source: A Survey on Machine Reading Comprehension: Tasks, Evaluation Metrics and Benchmark Datasets

Papers

Showing 251–300 of 1760 papers

TitleStatusHype
Annotating Entailment Relations for Shortanswer Questions—0
Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change—0
Benefits of Intermediate Annotations in Reading Comprehension—0
Annotating the MASC Corpus with BabelNet—0
A BERT based Sentiment Analysis and Key Entity Detection Approach for Online Financial Texts—0
CoMiC: Adapting a Short Answer Assessment System for Answer Selection—0
Annotating and Extracting Synthesis Process of All-Solid-State Batteries from Scientific Literature—0
Recent Advances in Multi-Choice Machine Reading Comprehension: A Survey on Methods and Datasets—0
AWS CORD-19 Search: A Neural Search Engine for COVID-19 Literature—0
An NLP-based Reading Tool for Aiding Non-native English Readers—0
A Framework for Learning Assessment through Multimodal Analysis of Reading Behaviour and Language Comprehension—0
Combining Formal and Distributional Models of Temporal and Intensional Semantics—0
A Vietnamese Dataset for Evaluating Machine Reading Comprehension—0
An MRC Framework for Semantic Role Labeling—0
A Vietnamese Dataset for Evaluating Machine Reading Comprehension—0
Automating Reading Comprehension by Generating Question and Answer Pairs—0
A Corpus of Text Data and Gaze Fixations from Autistic and Non-Autistic Adults—0
Combining Probabilistic Logic and Deep Learning for Self-Supervised Learning—0
Commonsense Evidence Generation and Injection in Reading Comprehension—0
Automating Idea Unit Segmentation and Alignment for Assessing Reading Comprehension via Summary Protocol Analysis—0
Automatic Word Segmentation and Part-of-Speech Tagging of Ancient Chinese Based on BERT Model—0
An Intelligent Recommendation-cum-Reminder System—0
Automatic True/False Question Generation for Educational Purpose—0
An Initial Investigation of Non-Native Spoken Question-Answering—0
A Framework and Dataset for Abstract Art Generation via CalligraphyGAN—0
Coherent Zero-Shot Visual Instruction Generation—0
Automatic Question Generation using Relative Pronouns and Adverbs—0
A Coordination-based Approach for Focused Learning in Knowledge-Based Systems—0
App-Aware Response Synthesis for User Reviews—0
Automatic Mining of Salient Events from Multiple Documents—0
Automatic learner summary assessment for reading comprehension—0
A Frame-based Sentence Representation for Machine Reading Comprehension—0
Automatic Judgment Prediction via Legal Reading Comprehension—0
Automatic Generation of Multiple-Choice Questions—0
中英文的文字蘊涵與閱讀測驗的初步探索 (An Exploration of Textual Entailment and Reading Comprehension for Chinese and English) [In Chinese]—0
An Exploration of Data Augmentation and Sampling Techniques for Domain-Agnostic Question Answering—0
Automatic Generation of Context-Based Fill-in-the-Blank Exercises Using Co-occurrence Likelihoods and Google n-grams—0
A Fine-grained Interpretability Evaluation Benchmark for Neural NLP—0
Coarse-to-Fine Question Answering for Long Documents—0
An Experimental Study of Deep Neural Network Models for Vietnamese Multiple-Choice Reading Comprehension—0
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs—0
Affective Common Sense Knowledge Acquisition for Sentiment Analysis—0
Automatic Evaluation vs. User Preference in Neural Textual QuestionAnswering over COVID-19 Scientific Literature—0
Automatic Entity State Annotation using the VerbNet Semantic Parser—0
A New Semantic Lexicon and Similarity Measure in Bangla—0
A Constituent-Centric Neural Architecture for Reading Comprehension—0
Co-Attention Hierarchical Network: Generating Coherent Long Distractors for Reading Comprehension—0
Collecting high-quality adversarial data for machine reading comprehension tasks with humans and models in the loop—0
Commonsense Inference in Natural Language Processing (COIN) - Shared Task Report—0
Automatic Classification of the Complexity of Nonfiction Texts in Portuguese for Early School Years—0
Show:102550
← PrevPage 6 of 36Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Rational Reasoner / IDOLTest80.6—Unverified
2AMR-LE-EnsembleTest80—Unverified
3MERIt(MERIt-deberta-v2-xxlarge )Test79.3—Unverified
4MERIt-deberta-v2-xxlarge deberta.v2.xxlarge.path.override_True.norm_1.1.0.w2.A100.cp200.s42Test79.3—Unverified
5Knowledge modelTest79.2—Unverified
6DeBERTa-v2-xxlarge-AMR-LE-ContrapositionTest77.2—Unverified
7LReasoner ensembleTest76.1—Unverified
8ELECTRA and ALBERTTest71—Unverified
9WWZTest69.7—Unverified
10xlnet-large-uncased [extended data]Test69.3—Unverified
#ModelMetricClaimedVerifiedStatus
1ALBERT (Ensemble)Accuracy91.4—Unverified
2Megatron-BERT (ensemble)Accuracy90.9—Unverified
3ALBERTxxlarge+DUMA(ensemble)Accuracy89.8—Unverified
4Megatron-BERTAccuracy89.5—Unverified
5XLNetAccuracy (Middle)88.6—Unverified
6DeBERTalargeAccuracy86.8—Unverified
7B10-10-10Accuracy85.7—Unverified
8RoBERTaAccuracy83.2—Unverified
9Orca 2-13BAccuracy82.87—Unverified
10Orca 2-7BAccuracy80.79—Unverified
#ModelMetricClaimedVerifiedStatus
1Golden TransformerAverage F10.94—Unverified
2MT5 LargeAverage F10.84—Unverified
3ruRoberta-large finetuneAverage F10.83—Unverified
4ruT5-large-finetuneAverage F10.82—Unverified
5Human BenchmarkAverage F10.81—Unverified
6ruT5-base-finetuneAverage F10.77—Unverified
7ruBert-large finetuneAverage F10.76—Unverified
8ruBert-base finetuneAverage F10.74—Unverified
9RuGPT3XL few-shotAverage F10.74—Unverified
10RuGPT3LargeAverage F10.73—Unverified
#ModelMetricClaimedVerifiedStatus
1RoBERTa-LargeOverall: F164.4—Unverified
2BERT-LargeOverall: F162.7—Unverified
3BiDAFOverall: F128.5—Unverified
#ModelMetricClaimedVerifiedStatus
1BERTMSE0.05—Unverified
#ModelMetricClaimedVerifiedStatus
1BERT pretrained on MIMIC-IIIAnswer F163.55—Unverified