SOTAVerified

Mathematical Problem-Solving

Papers

Showing 51–75 of 106 papers

TitleStatusHype
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning—0
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning—0
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models—0
How Do Large Language Monkeys Get Their Power (Laws)?—0
Improving Small-Scale Large Language Models Function Calling for Reasoning Tasks—0
JiuZhang 2.0: A Unified Chinese Pre-trained Language Model for Multi-task Mathematical Problem Solving—0
Kwai-STaR: Transform LLMs into State-Transition Reasoners—0
Large Language Models for Mathematical Reasoning: Progresses and Challenges—0
LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models—0
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems—0
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection—0
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task—0
PersonaMath: Enhancing Math Reasoning through Persona-Driven Data Augmentation—0
PoLAR: Polar-Decomposed Low-Rank Adapter Representation—0
Premise Order Matters in Reasoning with Large Language Models—0
Reasoning Models Can Be Effective Without Thinking—0
Scaling Autonomous Agents via Automatic Reward Modeling And Planning—0
Scaling Laws for Autoregressive Generative Modeling—0
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models—0
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models—0
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving—0
STRIVE: Structured Reasoning for Self-Improvement in Claim Verification—0
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving—0
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving—0
The Consensus Game: Language Model Generation via Equilibrium Search—0
Show:102550
← PrevPage 3 of 5Next →

No leaderboard results yet.