SOTAVerified

Math

Papers

Showing 901925 of 1596 papers

TitleStatusHype
Can Stories Help LLMs Reason? Curating Information Space Through Narrative0
ReasonAgain: Using Extractable Symbolic Programs to Evaluate Mathematical Reasoning0
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems0
Mixture of Parrots: Experts improve memorization more than reasoning0
MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning0
Polyak's Heavy Ball Method Achieves Accelerated Local Rate of Convergence under Polyak-Lojasiewicz Inequality0
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation0
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration0
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation0
No more hard prompts: SoftSRV prompting for synthetic data generation0
PromptHive: Bringing Subject Matter Experts Back to the Forefront with Collaborative Prompt Engineering for Educational Content Creation0
On Designing Effective RL Reward at Training Time for LLM Reasoning0
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology0
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems0
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning0
Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens0
SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented GenerationCode0
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math ReasoningCode0
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems0
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling0
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs0
Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps0
One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning TasksCode0
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning0
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning0
Show:102550
← PrevPage 37 of 64Next →

No leaderboard results yet.