SOTAVerified

Logical Reasoning

Papers

Showing 551–600 of 747 papers

TitleStatusHype
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models—0
Unlocking Temporal Question Answering for Large Language Models with Tailor-Made Reasoning LogicCode0
Deduction under Perturbed Evidence: Probing Student Simulation Capabilities of Large Language Models—0
Exploring Self-supervised Logic-enhanced Training for Large Language ModelsCode0
Query Structure Modeling for Inductive Logical Reasoning Over Knowledge GraphsCode0
Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization—0
Teaching Probabilistic Logical Reasoning to TransformersCode0
Atomic Inference for NLI with Generated Facts as AtomsCode0
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs—0
A Simple Generative Model of Logical Reasoning and Statistical Learning—0
Knowledge Authoring for Rules and Actions—0
Scalable Coupling of Deep Learning with Logical ReasoningCode0
Tackling Universal Properties of Minimal Trap Spaces of Boolean NetworksCode0
A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex TextCode0
The Dark Side of Explanations: Poisoning Recommender Systems with Counterfactual Examples—0
Sequential Recommendation with Probabilistic Logical ReasoningCode0
ChatABL: Abductive Learning via Natural Language Interaction with ChatGPT—0
LeafAI: query generator for clinical cohort discovery rivaling a human programmer—0
Scallop: A Language for Neurosymbolic Programming—0
Deep Manifold Learning for Reading Comprehension and Logical Reasoning Tasks with Polytuplet LossCode0
BloombergGPT: A Large Language Model for Finance—0
Logical Reasoning over Natural Language as Knowledge Representation: A SurveyCode0
Weakly Supervised Knowledge Transfer with Probabilistic Logical Reasoning for Object DetectionCode0
Attribution-Scores and Causal Counterfactuals as Explanations in Artificial Intelligence—0
A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT—0
Double Equivariance for Inductive Link Prediction for Both New Nodes and New Relation TypesCode0
Unifying Structure Reasoning and Language Model Pre-training for Complex Reasoning—0
A separation logic for sequences in pointer programs and its decidability—0
CogReact: A Reinforced Framework to Model Human Cognitive Reaction Modulated by Dynamic Intervention—0
LAMBADA: Backward Chaining for Automated Reasoning in Natural Language—0
APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical Reasoning—0
Towards High-Order Complementary Recommendation via Logical Reasoning NetworkCode0
Weisfeiler and Leman Go RelationalCode0
Neuro-Symbolic Spatio-Temporal Reasoning—0
Logical Tasks for Measuring Extrapolation and Rule ComprehensionCode0
Evident: a Development Methodology and a Knowledge Base Topology for Data Mining, Machine Learning and General Knowledge Management—0
Zero-Shot Classification by Logical Reasoning on Natural Language ExplanationsCode0
GammaE: Gamma Embeddings for Logical Queries on Knowledge GraphsCode0
TAPE: Assessing Few-shot Russian Language UnderstandingCode0
MetaLogic: Logical Reasoning Explanations with Fine-Grained StructureCode0
Investigating the Robustness of Natural Language Generation from Logical Forms via Counterfactual SamplesCode0
Inductive Logical Query Answering in Knowledge GraphsCode0
Join-Chain Network: A Logical Reasoning View of the Multi-head Attention in Transformer—0
To What Extent Do Natural Language Understanding Datasets Correlate to Logical Reasoning? A Method for Diagnosing Logical Reasoning.—0
Document-level Biomedical Relation Extraction Based on Multi-Dimensional Fusion Information and Multi-Granularity Logical ReasoningCode0
Type-dependent Prompt CycleQAG : Cycle Consistency for Multi-hop Question Generation—0
Towards Human-Compatible XAI: Explaining Data Differentials with Concept Induction over Background Knowledge—0
Time-aware Self-Attention Meets Logic Reasoning in Recommender Systems—0
Knowledge-based and Data-driven Reasoning and Learning for Ad Hoc Teamwork—0
A Scalable, Interpretable, Verifiable & Differentiable Logic Gate Convolutional Neural Network Architecture From Truth Tables—0
Show:102550
← PrevPage 12 of 15Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Claude OpusDelta_NoContext28.8—Unverified
2GPT-4oDelta_NoContext25.1—Unverified
3Gemini 1.5 ProDelta_NoContext23.4—Unverified
4GPT-4Delta_NoContext21.5—Unverified
5Command R+Delta_NoContext11.6—Unverified
6GPT-3.5Delta_NoContext11.2—Unverified
7Mixtral 8x7BDelta_NoContext6.4—Unverified
8Llama 3 8BDelta_NoContext4.9—Unverified
9Llama 3 70BDelta_NoContext2.9—Unverified
10Gemma 7BDelta_NoContext2.2—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, Direct)Accuracy64.8—Unverified
2PaLM 2 (few-shot, k=3, CoT)Accuracy57.2—Unverified
3OPT 66B (few-shot, k=3)Accuracy54—Unverified
4PaLM 540B (few-shot, k=3)Accuracy53.6—Unverified
5GPT-NeoX 20B (few-shot, k=3)Accuracy52.8—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy52.8—Unverified
7Chinchilla-70B (few-shot, k=5)Accuracy52.1—Unverified
8Bloomberg GPT 50B (few-shot, k=3)Accuracy50.8—Unverified
9Gopher-280B (few-shot, k=5)Accuracy50.7—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy84.9—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy65.8—Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy48.7—Unverified
4PaLM 540B (few-shot, k=3)Accuracy44.5—Unverified
5Gopher-280B (few-shot, k=5)Accuracy40.6—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy40.41—Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy37.67—Unverified
8GPT-NeoX (few-shot, k=3)Accuracy33.56—Unverified
9OPT 66B (few-shot, k=3)Accuracy28.08—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy91.2—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy61.2—Unverified
3Chinchilla-70B (few-shot, k=5)Accuracy59.7—Unverified
4Gopher-280B (few-shot, k=5)Accuracy49.2—Unverified
5PaLM 540B (few-shot, k=3)Accuracy38—Unverified
6BLOOM 176B (few-shot, k=3)Accuracy36.8—Unverified
7Bloomberg GPT (few-shot, k=3)Accuracy34.8—Unverified
8OPT 66B (few-shot, k=3)Accuracy31.2—Unverified
9GPT-NeoX (few-shot, k=3)Accuracy26—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy100—Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy96.4—Unverified
3PaLM 540B (few-shot, k=3)Accuracy39.6—Unverified
4BLOOM 176B (few-shot, k=3)Accuracy36.8—Unverified
5Chinchilla-70B (few-shot, k=5)Accuracy32—Unverified
6Bloomberg GPT (few-shot, k=3)Accuracy29.2—Unverified
7OPT 66B (few-shot, k=3)Accuracy23.6—Unverified
8GPT-NeoX (few-shot, k=3)Accuracy21.2—Unverified
9Gopher-280B (few-shot, k=5)Accuracy19—Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy44—Unverified
2PaLM-540B (few-shot, k=5)Accuracy42.4—Unverified
3PaLM-62B (few-shot, k=5)Accuracy36.5—Unverified
4Gopher-280B (few-shot, k=5)Accuracy35.1—Unverified
#ModelMetricClaimedVerifiedStatus
1PaLM-540B (few-shot, k=5)Accuracy73.9—Unverified
2Chinchilla-70B (few-shot, k=5)Accuracy68.3—Unverified
3PaLM-62B (few-shot, k=5)Accuracy65.4—Unverified
4Gopher-280B (few-shot, k=5)Accuracy61—Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy 83.7—Unverified
2RuGPT-3 LargeAccuracy 40.7—Unverified
3RuGPT-3 MediumAccuracy 38—Unverified
4RuGPT-3 SmallAccuracy 34—Unverified
#ModelMetricClaimedVerifiedStatus
1Human benchmarkAccuracy87—Unverified
2RuGPT-3 SmallAccuracy57.9—Unverified
3RuGPT-3 MediumAccuracy57.2—Unverified
4RuGPT-3 LargeAccuracy55.5—Unverified
#ModelMetricClaimedVerifiedStatus
1Chinchilla-70B (few-shot, k=5)Accuracy72.1—Unverified
2Gopher-280B (few-shot, k=5)Accuracy58.9—Unverified