SOTAVerified

Common Sense Reasoning

Common sense reasoning tasks are intended to require the model to go beyond pattern recognition. Instead, the model should use "common sense" or world knowledge to make inferences.

Papers

Showing 151–200 of 939 papers

TitleStatusHype
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive PrinciplesCode1
P-TA: Using Proximal Policy Optimization to Enhance Tabular Data Augmentation via Large Language Models—0
Mixture-of-Subspaces in Low-Rank AdaptationCode0
A Survey of Video Datasets for Grounded Event UnderstandingCode0
LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions—0
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented GenerationCode1
BAMO at SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common SenseCode0
Think out Loud: Emotion Deducing Explanation in Dialogues—0
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation—0
Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models—0
Generative AI-in-the-loop: Integrating LLMs and GPTs into the Next Generation Networks—0
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans—0
Every Answer Matters: Evaluating Commonsense with Probabilistic MeasuresCode0
Do Language Models Understand Morality? Towards a Robust Detection of Moral ContentCode0
Large Language Models as Evaluators for Recommendation ExplanationsCode1
RAG-based Crowdsourcing Task Decomposition via Masked Contrastive Learning with Prompts—0
ACCORD: Closing the Commonsense Measurability GapCode0
Extended Mind TransformersCode2
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language ModelsCode4
Easy Problems That LLMs Get WrongCode2
Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems—0
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search—0
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge—0
iREL at SemEval-2024 Task 9: Improving Conventional Prompting Methods for Brain TeasersCode0
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsCode2
Regressor-free Molecule Generation to Support Drug Response Prediction—0
Large Language Models are Effective Priors for Causal Graph Discovery—0
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question Answering—0
DaVinci at SemEval-2024 Task 9: Few-shot prompting GPT-3.5 for Unconventional Reasoning—0
Meta-Control: Automatic Model-based Control Synthesis for Heterogeneous Robot Skills—0
OpenBA-V2: Reaching 77.3% High Compression Ratio with Fast Multi-Stage PruningCode1
Soft Label PU Learning—0
The Power of Question Translation Training in Multilingual Reasoning: Broadened Scope and Deepened Insights—0
Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G—0
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models—0
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning—0
SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense—0
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of ExpertsCode3
Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution GapCode1
Concept Induction using LLMs: a user experiment for assessment—0
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences—0
Memory Sharing for Large Language Model based AgentsCode1
VLLMs Provide Better Context for Emotion Understanding Through Common Sense ReasoningCode1
Deep Reinforcement Learning-Based Approach for a Single Vehicle Persistent Surveillance Problem with Fuel Constraints—0
DELTA: Decomposed Efficient Long-Term Robot Task Planning using Large Language Models—0
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge GraphCode0
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach—0
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs—0
AILS-NTUA at SemEval-2024 Task 9: Cracking Brain Teasers: Transformer Models for Lateral Thinking PuzzlesCode0
ITCMA: A Generative Agent Based on a Computational Consciousness Structure—0
Show:102550
← PrevPage 4 of 19Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1ST-MoE-32B 269B (fine-tuned)Accuracy96.1—Unverified
2Unicorn 11B (fine-tuned)Accuracy91.3—Unverified
3CompassMTL 567M with TailorAccuracy90.5—Unverified
4CompassMTL 567MAccuracy89.6—Unverified
5UnifiedQA 11B (fine-tuned)Accuracy89.4—Unverified
6Claude 3 Opus (5-shot)Accuracy88.5—Unverified
7GPT-4 (5-shot)Accuracy87.5—Unverified
8ExDeBERTa 567MAccuracy87—Unverified
9LLaMA-2 13B + MixLoRAAccuracy86.3—Unverified
10LLaMA3 8B+MoSLoRAAccuracy85.8—Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4 (few-shot, k=25)Accuracy96.4—Unverified
2PaLM 2 (few-shot, CoT, SC)Accuracy95.1—Unverified
3Shivaay (4B, few-shot, k=8)Accuracy91.04—Unverified
4StupidLLMAccuracy91.03—Unverified
5Claude 2 (few-shot, k=5)Accuracy91—Unverified
6Claude 1.3 (few-shot, k=5)Accuracy90—Unverified
7PaLM 540B (Self Improvement, Self Consistency)Accuracy89.8—Unverified
8PaLM 540B (Self Consistency)Accuracy88.7—Unverified
9PaLM 540B (Self Improvement, CoT Prompting)Accuracy88.3—Unverified
10PaLM 540B (Self Improvement, Standard-Prompting)Accuracy87.2—Unverified
#ModelMetricClaimedVerifiedStatus
1ST-MoE-32B 269B (fine-tuned)Accuracy95.2—Unverified
2LLaMA 3 8B+MoSLoRA (fine-tuned)Accuracy90.5—Unverified
3PaLM 2-L (1-shot)Accuracy89.7—Unverified
4PaLM 2-M (1-shot)Accuracy88—Unverified
5LLaMA-3 8B + MixLoRAAccuracy86.5—Unverified
6Camelidae-8×34BAccuracy86.2—Unverified
7PaLM 2-S (1-shot)Accuracy85.6—Unverified
8LLaMA 65B + CFG (0-shot)Accuracy84.2—Unverified
9GAL 120B (0-shot)Accuracy83.8—Unverified
10LLaMA-2 13B + MixLoRAAccuracy83.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Turing NLR v5 XXL 5.4B (fine-tuned)EM95.9—Unverified
2ST-MoE-32B 269B (fine-tuned)EM95.1—Unverified
3T5-11BF194.1—Unverified
4DeBERTa-1.5BEM94.1—Unverified
5PaLM 540B (finetuned)EM94—Unverified
6Vega v2 6B (fine-tuned)EM93.9—Unverified
7PaLM 2-L (one-shot)F193.8—Unverified
8T5-XXL 11B (fine-tuned)EM93.4—Unverified
9PaLM 2-M (one-shot)F192.4—Unverified
10PaLM 2-S (one-shot)F192.1—Unverified