SOTAVerified

Logical Reasoning

Papers

Showing 101150 of 747 papers

TitleStatusHype
Neuro-Symbolic Integration Brings Causal and Reliable Reasoning ProofsCode1
Chain of Images for Intuitively ReasoningCode1
Dynamics of Instruction Tuning: Each Ability of Large Language Models Has Its Own Growth PaceCode1
DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to DeterminacyCode1
Plan, Verify and Switch: Integrated Reasoning with Diverse X-of-ThoughtsCode1
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversCode1
Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real WorldCode1
GLoRE: Evaluating Logical Reasoning of Large Language ModelsCode1
Improving Large Language Models in Event Relation Logical PredictionCode1
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical ReasoningCode1
Do PLMs Know and Understand Ontological Knowledge?Code1
HAE-RAE Bench: Evaluation of Korean Knowledge in Language ModelsCode1
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking PuzzlesCode1
Evolving Scientific Discovery by Unifying Data and Background Knowledge with AI HilbertCode1
Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicCode1
COLLIE: Systematic Construction of Constrained Text Generation TasksCode1
EFO_k-CQA: Towards Knowledge Graph Complex Query Answering beyond Set OperationCode1
IDOL: Indicator-oriented Logic Pre-training for Logical ReasoningCode1
Modeling Hierarchical Reasoning Chains by Linking Discourse Units and Key Phrases for Reading ComprehensionCode1
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and BeyondCode1
Certified Deductive Reasoning with Language ModelsCode1
Deductive Verification of Chain-of-Thought ReasoningCode1
Counterfactual reasoning: Testing language models' understanding of hypothetical scenariosCode1
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical ReasoningCode1
LogiCoT: Logical Chain-of-Thought Instruction-TuningCode1
Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought PromptingCode1
Wasserstein-Fisher-Rao Embedding: Logical Query Embeddings with Local Comparison and Global TransportCode1
Improved Logical Reasoning of Language Models via Differentiable Symbolic ProgrammingCode1
Complex Logical Reasoning over Knowledge Graphs using Large Language ModelsCode1
Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4Code1
Explicit Planning Helps Language Models in Logical ReasoningCode1
Neural Graph Reasoning: Complex Logical Query Answering Meets Graph DatabasesCode1
Natural Language Reasoning, A SurveyCode1
Domain Specific Question Answering Over Knowledge Graphs Using Logical Programming and Large Language ModelsCode1
ChatCAD: Interactive Computer-Aided Diagnosis on Medical Image using Large Language ModelsCode1
A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and InteractivityCode1
Logical Message Passing Networks with One-hop Inference on Atomic FormulasCode1
Mind Reasoning Manners: Enhancing Type Perception for Generalized Zero-shot Logical Reasoning over TextCode1
Large Language Models are Better Reasoners with Self-VerificationCode1
On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot ReasoningCode1
Counterfactual reasoning: Do language models need world knowledge for causal understanding?Code1
UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical ExpressionCode1
NQE: N-ary Query Embedding for Complex Query Answering over Hyper-Relational Knowledge GraphsCode1
Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningCode1
Neural Methods for Logical Reasoning Over Knowledge GraphsCode1
FOLIO: Natural Language Reasoning with First-Order LogicCode1
Semantic Probabilistic Layers for Neuro-Symbolic LearningCode1
TFLEX: Temporal Feature-Logic Embedding Framework for Complex Reasoning over Temporal Knowledge GraphCode1
On the Paradox of Learning to Reason from DataCode1
Logiformer: A Two-Branch Graph Transformer Network for Interpretable Logical ReasoningCode1
Show:102550
← PrevPage 3 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Claude OpusDelta_NoContext28.8Unverified
2GPT-4oDelta_NoContext25.1Unverified
3Gemini 1.5 ProDelta_NoContext23.4Unverified
4GPT-4Delta_NoContext21.5Unverified
5Command R+Delta_NoContext11.6Unverified
6GPT-3.5Delta_NoContext11.2Unverified
7Mixtral 8x7BDelta_NoContext6.4Unverified
8Llama 3 8BDelta_NoContext4.9Unverified
9Llama 3 70BDelta_NoContext2.9Unverified
10Gemma 7BDelta_NoContext2.2Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, Direct)Accuracy64.8Unverified
2PaLM 2 (few-shot, k=3, CoT)Accuracy57.2Unverified
3OPT 66B (few-shot, k=3)Accuracy54Unverified
4PaLM 540B (few-shot, k=3)Accuracy53.6Unverified
5GPT-NeoX 20B (few-shot, k=3)Accuracy52.8Unverified
6BLOOM 176B (few-shot, k=3)Accuracy52.8Unverified
7Chinchilla-70B (few-shot, k=5)Accuracy52.1Unverified
8Bloomberg GPT 50B (few-shot, k=3)Accuracy50.8Unverified
9Gopher-280B (few-shot, k=5)Accuracy50.7Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy84.9Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy65.8Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy48.7Unverified
4PaLM 540B (few-shot, k=3)Accuracy44.5Unverified
5Gopher-280B (few-shot, k=5)Accuracy40.6Unverified
6BLOOM 176B (few-shot, k=3)Accuracy40.41Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy37.67Unverified
8GPT-NeoX (few-shot, k=3)Accuracy33.56Unverified
9OPT 66B (few-shot, k=3)Accuracy28.08Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy91.2Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy61.2Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy59.7Unverified
4Gopher-280B (few-shot, k=5)Accuracy49.2Unverified
5PaLM 540B (few-shot, k=3)Accuracy38Unverified
6BLOOM 176B (few-shot, k=3)Accuracy36.8Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy34.8Unverified
8OPT 66B (few-shot, k=3)Accuracy31.2Unverified
9GPT-NeoX (few-shot, k=3)Accuracy26Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy100Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy96.4Unverified
3PaLM 540B (few-shot, k=3)Accuracy39.6Unverified
4BLOOM 176B (few-shot, k=3)Accuracy36.8Unverified
5Chinchilla-70B (few-shot, k=5)Accuracy32Unverified
6Bloomberg GPT (few-shot, k=3)Accuracy29.2Unverified
7OPT 66B (few-shot, k=3)Accuracy23.6Unverified
8GPT-NeoX (few-shot, k=3)Accuracy21.2Unverified
9Gopher-280B (few-shot, k=5)Accuracy19Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy44Unverified
2PaLM-540B (few-shot, k=5)Accuracy42.4Unverified
3PaLM-62B (few-shot, k=5)Accuracy36.5Unverified
4Gopher-280B (few-shot, k=5)Accuracy35.1Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM-540B (few-shot, k=5)Accuracy73.9Unverified
2Chinchilla-70B (few-shot, k=5)Accuracy68.3Unverified
3PaLM-62B (few-shot, k=5)Accuracy65.4Unverified
4Gopher-280B (few-shot, k=5)Accuracy61Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy 83.7Unverified
2RuGPT-3 LargeAccuracy 40.7Unverified
3RuGPT-3 MediumAccuracy 38Unverified
4RuGPT-3 SmallAccuracy 34Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy87Unverified
2RuGPT-3 SmallAccuracy57.9Unverified
3RuGPT-3 MediumAccuracy57.2Unverified
4RuGPT-3 LargeAccuracy55.5Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy72.1Unverified
2Gopher-280B (few-shot, k=5)Accuracy58.9Unverified