SOTAVerified

GSM8K

Papers

Showing 251–275 of 439 papers

TitleStatusHype
A Careful Examination of Large Language Model Performance on Grade School Arithmetic—0
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning—0
Uncertainty Aware Learning for Language Model Alignment—0
No Train Still Gain. Unleash Mathematical Reasoning of Large Language Models with Monte Carlo Tree Search Guided by Energy Function—0
Nudging: Inference-time Alignment of LLMs via Guided Decoding—0
Fine-Grained Self-Endorsement Improves Factuality and Reasoning—0
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning—0
On Designing Effective RL Reward at Training Time for LLM Reasoning—0
Uncertainty-Aware Search and Value Models: Mitigating Search Scaling Flaws in LLMs—0
Making Large Language Models Better Reasoners with Step-Aware Verifier—0
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty—0
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation—0
Orca-Math: Unlocking the potential of SLMs in Grade School Math—0
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree—0
Exploring an LM to generate Prolog Predicates from Mathematics Questions—0
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning—0
Explicit Knowledge Transfer for Weakly-Supervised Code Generation—0
PARAMANU-GANITA: Language Model with Mathematical Capabilities—0
Patience Is The Key to Large Language Model Reasoning—0
PersonaMath: Enhancing Math Reasoning through Persona-Driven Data Augmentation—0
Pheromone-based Learning of Optimal Reasoning Paths—0
Excessive Reasoning Attack on Reasoning LLMs—0
Evolving LLMs' Self-Refinement Capability via Iterative Preference Optimization—0
Plan for Speed -- Dilated Scheduling for Masked Diffusion Language Models—0
PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models—0
Show:102550
← PrevPage 11 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1XolverAccuracy98.1—Unverified
2Orange-mini0-shot MRR98—Unverified