SOTAVerified

Common Sense Reasoning

Common sense reasoning tasks are intended to require the model to go beyond pattern recognition. Instead, the model should use "common sense" or world knowledge to make inferences.

Papers

Showing 651–700 of 939 papers

TitleStatusHype
A Knowledge-Aware Sequence-to-Tree Network for Math Word Problem Solving—0
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks—0
Machine Reasoning: Technology, Dilemma and Future—0
Learning Physical Common Sense as Knowledge Graph Completion via BERT Data Augmentation and Constrained Tucker Factorization—0
Dutch Humor Detection by Generating Negative Examples—0
GO FIGURE: A Meta Evaluation of Factuality in Summarization—0
Thinking Fast and Slow in AI—0
Do Language Embeddings Capture Scales?—0
Hierarchical Relational Inference—0
Creative Captioning: An AI Grand Challenge Based on the Dixit Board Game—0
Zero-Shot Learning with Common Sense Knowledge Graphs—0
Multi-modal Cooking Workflow Construction for Food Recipes—0
Commonsense Knowledge in Wikidata—0
Learning Object Placement by Inpainting for Compositional Data Augmentation—0
CS-NET at SemEval-2020 Task 4: Siamese BERT for ComVECode0
Understanding Spatial Relations through Multiple Modalities—0
Pasadena: Perceptually Aware and Stealthy Adversarial Denoise Attack—0
Robustness to Spurious Correlations via Human AnnotationsCode0
Explainable Inference on Sequential Data via Memory-TrackingCode0
LMVE at SemEval-2020 Task 4: Commonsense Validation and Explanation using Pretraining Language Model—0
Machine Common Sense—0
CUHK at SemEval-2020 Task 4: CommonSense Explanation, Reasoning and Prediction with Multi-task Learning—0
Consolidating Commonsense Knowledge—0
Language Models as Fact Checkers?—0
Analogical Proportions—0
Fractional trends and cycles in macroeconomic time series—0
Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language Models—0
Temporal Common Sense Acquisition with Minimal Supervision—0
The Sensitivity of Language Models and Humans to Winograd Schema PerturbationsCode0
The ILASP system for Inductive Learning of Answer Set Programs—0
Mandarinograd: A Chinese Collection of Winograd Schemas—0
Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense—0
Ecological Semantics: Programming Environments for Situated Language Understanding—0
1D Probabilistic Undersampling Pattern Optimization for MR Image ReconstructionCode0
Active Model Estimation in Markov Decision Processes—0
Learning-based Practical Smartphone Eavesdropping with Built-in Accelerometer—0
KoGuN: Accelerating Deep Reinforcement Learning via Integrating Human Suboptimal Knowledge—0
A Machine Consciousness architecture based on Deep Learning and Gaussian Processes—0
Debate Dynamics for Human-comprehensible Fact-checking on Knowledge Graphs—0
Using ConceptNet to Teach Common Sense to an Automated Theorem Prover—0
A Logical Model for Supporting Social Commonsense Knowledge Acquisition—0
Design and Implementation of Linked Planning Domain Definition Language—0
That and There: Judging the Intent of Pointing Actions with Robotic ArmsCode0
Generating Interactive Worlds with Text—0
CommonGen: A Constrained Text Generation Challenge for Generative Commonsense ReasoningCode0
Why Do Masked Neural Language Models Still Need Common Sense Knowledge?—0
KARNA at COIN Shared Task 1: Bidirectional Encoder Representations from Transformers with relational knowledge for machine comprehension with common sense—0
Commonsense about Human Senses: Labeled Data Collection Processes—0
How Pre-trained Word Representations Capture Commonsense Physical Comparisons—0
Pingan Smart Health and SJTU at COIN - Shared Task: utilizing Pre-trained Language Models and Common-sense Knowledge in Machine Reading Tasks—0
Show:102550
← PrevPage 14 of 19Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ST-MoE-32B 269B (fine-tuned)Accuracy96.1—Unverified
2Unicorn 11B (fine-tuned)Accuracy91.3—Unverified
3CompassMTL 567M with TailorAccuracy90.5—Unverified
4CompassMTL 567MAccuracy89.6—Unverified
5UnifiedQA 11B (fine-tuned)Accuracy89.4—Unverified
6Claude 3 Opus (5-shot)Accuracy88.5—Unverified
7GPT-4 (5-shot)Accuracy87.5—Unverified
8ExDeBERTa 567MAccuracy87—Unverified
9LLaMA-2 13B + MixLoRAAccuracy86.3—Unverified
10LLaMA3 8B+MoSLoRAAccuracy85.8—Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4 (few-shot, k=25)Accuracy96.4—Unverified
2PaLM 2 (few-shot, CoT, SC)Accuracy95.1—Unverified
3Shivaay (4B, few-shot, k=8)Accuracy91.04—Unverified
4StupidLLMAccuracy91.03—Unverified
5Claude 2 (few-shot, k=5)Accuracy91—Unverified
6Claude 1.3 (few-shot, k=5)Accuracy90—Unverified
7PaLM 540B (Self Improvement, Self Consistency)Accuracy89.8—Unverified
8PaLM 540B (Self Consistency)Accuracy88.7—Unverified
9PaLM 540B (Self Improvement, CoT Prompting)Accuracy88.3—Unverified
10PaLM 540B (Self Improvement, Standard-Prompting)Accuracy87.2—Unverified
#ModelMetricClaimedVerifiedStatus
1ST-MoE-32B 269B (fine-tuned)Accuracy95.2—Unverified
2LLaMA 3 8B+MoSLoRA (fine-tuned)Accuracy90.5—Unverified
3PaLM 2-L (1-shot)Accuracy89.7—Unverified
4PaLM 2-M (1-shot)Accuracy88—Unverified
5LLaMA-3 8B + MixLoRAAccuracy86.5—Unverified
6Camelidae-8×34BAccuracy86.2—Unverified
7PaLM 2-S (1-shot)Accuracy85.6—Unverified
8LLaMA 65B + CFG (0-shot)Accuracy84.2—Unverified
9GAL 120B (0-shot)Accuracy83.8—Unverified
10LLaMA-2 13B + MixLoRAAccuracy83.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Turing NLR v5 XXL 5.4B (fine-tuned)EM95.9—Unverified
2ST-MoE-32B 269B (fine-tuned)EM95.1—Unverified
3T5-11BF194.1—Unverified
4DeBERTa-1.5BEM94.1—Unverified
5PaLM 540B (finetuned)EM94—Unverified
6Vega v2 6B (fine-tuned)EM93.9—Unverified
7PaLM 2-L (one-shot)F193.8—Unverified
8T5-XXL 11B (fine-tuned)EM93.4—Unverified
9PaLM 2-M (one-shot)F192.4—Unverified
10PaLM 2-S (one-shot)F192.1—Unverified