SOTAVerified

Math

Papers

Showing 201250 of 1596 papers

TitleStatusHype
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal ModelsCode2
A Survey of Deep Learning for Mathematical ReasoningCode2
Dynamic Early Exit in Reasoning ModelsCode2
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought ReasoningCode2
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsCode2
ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique PipelineCode2
Meta Prompting for AI SystemsCode2
GPT Can Solve Mathematical Problems Without a CalculatorCode2
Meta-Design Matters: A Self-Design Multi-Agent SystemCode2
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsCode2
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language ModelsCode2
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in ChineseCode2
Measuring Mathematical Problem Solving With the MATH DatasetCode2
Measuring Multimodal Mathematical Reasoning with MATH-Vision DatasetCode2
Full Page Handwriting Recognition via Image to Sequence ExtractionCode2
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning GapCode2
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math DataCode2
A Comparative Study on Reasoning Patterns of OpenAI's o1 ModelCode2
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language ModelsCode2
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsCode2
MegaMath: Pushing the Limits of Open Math CorporaCode2
Balancing LoRA Performance and Efficiency with Simple Shard SharingCode2
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics BenchmarkCode2
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical CodeCode2
MAS-Zero: Designing Multi-Agent Systems with Zero SupervisionCode2
Flaming-hot Initiation with Regular Execution Sampling for Large Language ModelsCode2
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical ReasoningCode2
Archon: An Architecture Search Framework for Inference-Time TechniquesCode2
AbstentionBench: Reasoning LLMs Fail on Unanswerable QuestionsCode2
Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPOCode2
MathPile: A Billion-Token-Scale Pretraining Corpus for MathCode2
Memorizing TransformersCode2
On the Emergence of Thinking in LLMs I: Searching for the Right IntuitionCode2
SciInstruct: a Self-Reflective Instruction Annotated Dataset for Training Scientific Language ModelsCode2
Expression Syntax Information Bottleneck for Math Word ProblemsCode1
M1: Towards Scalable Test-Time Compute with Mamba Reasoning ModelsCode1
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMsCode1
A Probabilistic Inference Approach to Inference-Time Scaling of LLMs using Particle-Based Monte Carlo MethodsCode1
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?Code1
Explaining Datasets in Words: Statistical Models with Natural Language ParametersCode1
Can an AI Win Ghana's National Science and Maths Quiz? An AI Grand Challenge for EducationCode1
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement LearningCode1
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy PreservationCode1
Evolving Prompts In-Context: An Open-ended, Self-replicating PerspectiveCode1
LLMThinkBench: Towards Basic Math Reasoning and Overthinking in Large Language ModelsCode1
EXAONE Deep: Reasoning Enhanced Language ModelsCode1
LoRA Soups: Merging LoRAs for Practical Skill Composition TasksCode1
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language ModelsCode1
Building Dataset for Grounding of Formulae — Annotating Coreference Relations Among Math IdentifiersCode1
Broken Neural Scaling LawsCode1
Show:102550
← PrevPage 5 of 32Next →

No leaderboard results yet.