SOTAVerified

Common Sense Reasoning

Common sense reasoning tasks are intended to require the model to go beyond pattern recognition. Instead, the model should use "common sense" or world knowledge to make inferences.

Papers

Showing 1–10 of 939 papers

TitleStatusHype
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes—0
LoSiA: Efficient High-Rank Fine-Tuning via Subnet Localization and OptimizationCode0
CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation—0
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits—0
Prime the search: Using large language models for guiding geometric task and motion planning by warm-starting tree searchCode0
AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentCode0
ATLAS: Learning to Optimally Memorize the Context at Test Time—0
Spatial Knowledge Graph-Guided Multimodal Synthesis—0
CaseEdit: Enhancing Localized Commonsense Reasoning via Null-Space Constrained Knowledge Editing in Small Parameter Language Models—0
Align-GRAG: Reasoning-Guided Dual Alignment for Graph Retrieval-Augmented Generation—0
Show:102550
← PrevPage 1 of 94Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 (few-shot, k=25)Accuracy96.4—Unverified
2PaLM 2 (few-shot, CoT, SC)Accuracy95.1—Unverified
3Shivaay (4B, few-shot, k=8)Accuracy91.04—Unverified
4StupidLLMAccuracy91.03—Unverified
5Claude 2 (few-shot, k=5)Accuracy91—Unverified
6Claude 1.3 (few-shot, k=5)Accuracy90—Unverified
7PaLM 540B (Self Improvement, Self Consistency)Accuracy89.8—Unverified
8PaLM 540B (Self Consistency)Accuracy88.7—Unverified
9PaLM 540B (Self Improvement, CoT Prompting)Accuracy88.3—Unverified
10PaLM 540B (Self Improvement, Standard-Prompting)Accuracy87.2—Unverified