SOTAVerified

GSM8K

Papers

Showing 151–175 of 439 papers

TitleStatusHype
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle SolvingCode1
Large Language Models as OptimizersCode1
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward PassesCode1
Over-Reasoning and Redundant Calculation of Large Language ModelsCode1
MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought ThinkingCode1
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided InterventionsCode1
Fine-Grained Self-Endorsement Improves Factuality and Reasoning—0
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning—0
CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs—0
Automatic Robustness Stress Testing of LLMs as Mathematical Problem Solvers—0
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty—0
A Careful Examination of Large Language Model Performance on Grade School Arithmetic—0
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree—0
Cool-Fusion: Fuse Large Language Models without Training—0
Automatic Prompt Selection for Large Language Models—0
ControlMath: Controllable Data Generation Promotes Math Generalist Models—0
Meaning-Typed Programming: Language Abstraction and Runtime for Model-Integrated Applications—0
Exploring an LM to generate Prolog Predicates from Mathematics Questions—0
Explicit Knowledge Transfer for Weakly-Supervised Code Generation—0
Contrastive Decoding Improves Reasoning in Large Language Models—0
Excessive Reasoning Attack on Reasoning LLMs—0
Evolving LLMs' Self-Refinement Capability via Iterative Preference Optimization—0
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost—0
Evolutionary Pre-Prompt Optimization for Mathematical Reasoning—0
Evaluation of LLMs for mathematical problem solving—0
Show:102550
← PrevPage 7 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1XolverAccuracy98.1—Unverified
2Orange-mini0-shot MRR98—Unverified