SOTAVerified

GSM8K

Papers

Showing 251–275 of 439 papers

TitleStatusHype
Iterative Reasoning Preference Optimization—0
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning—0
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?—0
Kwai-STaR: Transform LLMs into State-Transition Reasoners—0
KwaiYiiMath: Technical Report—0
Large Language Models as Analogical Reasoners—0
Large Language Models Can Self-Improve—0
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge—0
LearnAlign: Reasoning Data Selection for Reinforcement Learning in Large Language Models Based on Improved Gradient Alignment—0
Learning to Rank Chain-of-Thought: An Energy-Based Approach with Outcome Supervision—0
Learning to Reason via Self-Iterative Process Feedback for Small Language Models—0
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint—0
Let's Reinforce Step by Step—0
Let's reward step by step: Step-Level reward model as the Navigators for Reasoning—0
Leveraging Uncertainty Estimation for Efficient LLM Routing—0
LiteSearch: Efficacious Tree Search for LLM—0
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models—0
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ—0
Meaning-Typed Programming: Language Abstraction and Runtime for Model-Integrated Applications—0
DavIR: Data Selection via Implicit Reward for Large Language Models—0
Local Prompt Optimization—0
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems—0
Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models—0
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing—0
MALT: Improving Reasoning with Multi-Agent LLM Training—0
Show:102550
← PrevPage 11 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1XolverAccuracy98.1—Unverified
2Orange-mini0-shot MRR98—Unverified