SOTAVerified

Logical Reasoning

Papers

Showing 110 of 747 papers

TitleStatusHype
FEVO: Financial Knowledge Expansion and Reasoning Evolution for Large Language Models0
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning0
Discrete JEPA: Learning Discrete Token Representations without Reconstruction0
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making0
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language ModelsCode5
Motion-R1: Chain-of-Thought Reasoning and Reinforcement Learning for Human Motion Generation0
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving0
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games0
EviNet: Evidential Reasoning Network for Resilient Graph Learning in the Open and Noisy EnvironmentsCode0
Are LLMs Reliable Translators of Logical Reasoning Across Lexically Diversified Contexts?Code0
Show:102550
← PrevPage 1 of 75Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1PaLM 2 (few-shot, k=3, CoT)Accuracy100Unverified
2PaLM 2 (few-shot, k=3, Direct)Accuracy96.4Unverified
3PaLM 540B (few-shot, k=3)Accuracy39.6Unverified
4BLOOM 176B (few-shot, k=3)Accuracy36.8Unverified
5Chinchilla-70B (few-shot, k=5)Accuracy32Unverified
6Bloomberg GPT (few-shot, k=3)Accuracy29.2Unverified
7OPT 66B (few-shot, k=3)Accuracy23.6Unverified
8GPT-NeoX (few-shot, k=3)Accuracy21.2Unverified
9Gopher-280B (few-shot, k=5)Accuracy19Unverified