SOTAVerified

Logical Reasoning

Papers

Showing 601–650 of 747 papers

TitleStatusHype
Reduced Implication-bias Logic Loss for Neuro-Symbolic Learning—0
Language models show human-like content effects on reasoning tasksCode0
Emotion Recognition in Conversation using Probabilistic Soft Logic—0
Discourse-Aware Graph Networks for Textual Logical Reasoning—0
AnaLog: Testing Analytical and Deductive Logic Learnability in Language Models—0
Learning Symmetric Rules with SATNetCode0
Towards Unifying Perceptual Reasoning and Logical Reasoning—0
TAR: Neural Logical Reasoning across TBox and ABox—0
Reasoning over Logically Interacted Conditions for Question Answering—0
RobustLR: Evaluating Robustness to Logical Perturbation in Deductive ReasoningCode0
FLEX: Feature-Logic Embedding Framework for CompleX Knowledge Graph ReasoningCode0
Logical Reasoning with Span-Level Predictions for Interpretable and Robust NLI ModelsCode0
Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning—0
LogiGAN: Learning Logical Reasoning via Adversarial Pre-training—0
Graph Neural Networks for Propositional Model Counting—0
Table-based Fact Verification with Self-adaptive Mixture of ExpertsCode0
Reasoning with Multi-Structure Commonsense Knowledge in Visual Dialog—0
Enhancing Neural Mathematical Reasoning by Abductive Combination with Symbolic Library—0
A Densely Connected Criss-Cross Attention Network for Document-level Relation Extraction—0
A Neural-Symbolic Approach to Natural Language UnderstandingCode0
What Makes Reading Comprehension Questions Difficult?Code0
Towards Unifying Logical Entailment and Statistical Estimation—0
MUC-driven Feature Importance Measurement and Adversarial Analysis for Random Forest—0
JAMES: Normalizing Job Titles with Multi-Aspect Graph Embeddings and Reasoning—0
Logical Reasoning for Task Oriented Dialogue Systems—0
Neural Logic Analogy Learning—0
Reasoning Like Program Executors—0
Combining Commonsense Reasoning and Knowledge Acquisition to Guide Deep Learning in Robotics—0
BTPK-based interpretable method for NER tasks based on Talmudic Public Announcement Logic—0
Scales and Hedges in a Logic with Analogous Semantics—0
Emergent Symbols through Binding in External Memory—0
Quantifying Adaptability in Pre-trained Language Models with 500 Tasks—0
MANGO: Enhancing the Robustness of VQA Models via Adversarial Noise Generation—0
Can BERT Conduct Logical Reasoning? On the Difficulty of Learning to Reason from Data—0
FaiRR: Faithful and Robust Deductive Reasoning over Natural Language—0
Does Entity Abstraction Help Generative Transformers Reason?—0
Modeling Associative Reasoning Processes—0
Explainability Is in the Mind of the Beholder: Establishing the Foundations of Explainable Artificial Intelligence—0
Graph Collaborative Reasoning—0
The theory of quantitative trading—0
LoNLI: An Extensible Framework for Testing Diverse Logical Reasoning Capabilities for NLI—0
Scallop: From Probabilistic Deductive Databases to Scalable Differentiable Reasoning—0
Two-stage Rule-induction Visual Reasoning on RPMs with an Application to Video Prediction—0
What Makes Machine Reading Comprehension Questions Difficult? Investigating Variation in Passage Sources and Question Types—0
CausalR: Causal Reasoning over Natural Language Rulebases—0
Table-based Fact Verification with Self-adaptive Mixture of Experts—0
AbductionRules: Training Transformers to Explain Unexpected Inputs—0
Logic-Driven Context Extension and Data Augmentation for Logical Reasoning of Text—0
Reasoning Like Program Executors—0
Automated scholarly paper review: Concepts, technologies, and challenges—0
Show:102550
← PrevPage 13 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Claude OpusDelta_NoContext28.8—Unverified
2GPT-4oDelta_NoContext25.1—Unverified
3Gemini 1.5 ProDelta_NoContext23.4—Unverified
4GPT-4Delta_NoContext21.5—Unverified
5Command R+Delta_NoContext11.6—Unverified
6GPT-3.5Delta_NoContext11.2—Unverified
7Mixtral 8x7BDelta_NoContext6.4—Unverified
8Llama 3 8BDelta_NoContext4.9—Unverified
9Llama 3 70BDelta_NoContext2.9—Unverified
10Gemma 7BDelta_NoContext2.2—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, Direct)Accuracy64.8—Unverified
2PaLM 2 (few-shot, k=3, CoT)Accuracy57.2—Unverified
3OPT 66B (few-shot, k=3)Accuracy54—Unverified
4PaLM 540B (few-shot, k=3)Accuracy53.6—Unverified
5GPT-NeoX 20B (few-shot, k=3)Accuracy52.8—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy52.8—Unverified
7Chinchilla-70B (few-shot, k=5)Accuracy52.1—Unverified
8Bloomberg GPT 50B (few-shot, k=3)Accuracy50.8—Unverified
9Gopher-280B (few-shot, k=5)Accuracy50.7—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy84.9—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy65.8—Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy48.7—Unverified
4PaLM 540B (few-shot, k=3)Accuracy44.5—Unverified
5Gopher-280B (few-shot, k=5)Accuracy40.6—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy40.41—Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy37.67—Unverified
8GPT-NeoX (few-shot, k=3)Accuracy33.56—Unverified
9OPT 66B (few-shot, k=3)Accuracy28.08—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy91.2—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy61.2—Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy59.7—Unverified
4Gopher-280B (few-shot, k=5)Accuracy49.2—Unverified
5PaLM 540B (few-shot, k=3)Accuracy38—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy36.8—Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy34.8—Unverified
8OPT 66B (few-shot, k=3)Accuracy31.2—Unverified
9GPT-NeoX (few-shot, k=3)Accuracy26—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy100—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy96.4—Unverified
3PaLM 540B (few-shot, k=3)Accuracy39.6—Unverified
4BLOOM 176B (few-shot, k=3)Accuracy36.8—Unverified
5Chinchilla-70B (few-shot, k=5)Accuracy32—Unverified
6Bloomberg GPT (few-shot, k=3)Accuracy29.2—Unverified
7OPT 66B (few-shot, k=3)Accuracy23.6—Unverified
8GPT-NeoX (few-shot, k=3)Accuracy21.2—Unverified
9Gopher-280B (few-shot, k=5)Accuracy19—Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy44—Unverified
2PaLM-540B (few-shot, k=5)Accuracy42.4—Unverified
3PaLM-62B (few-shot, k=5)Accuracy36.5—Unverified
4Gopher-280B (few-shot, k=5)Accuracy35.1—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM-540B (few-shot, k=5)Accuracy73.9—Unverified
2Chinchilla-70B (few-shot, k=5)Accuracy68.3—Unverified
3PaLM-62B (few-shot, k=5)Accuracy65.4—Unverified
4Gopher-280B (few-shot, k=5)Accuracy61—Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy 83.7—Unverified
2RuGPT-3 LargeAccuracy 40.7—Unverified
3RuGPT-3 MediumAccuracy 38—Unverified
4RuGPT-3 SmallAccuracy 34—Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy87—Unverified
2RuGPT-3 SmallAccuracy57.9—Unverified
3RuGPT-3 MediumAccuracy57.2—Unverified
4RuGPT-3 LargeAccuracy55.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy72.1—Unverified
2Gopher-280B (few-shot, k=5)Accuracy58.9—Unverified