SOTAVerified

GSM8K

Papers

Showing 226–250 of 439 papers

TitleStatusHype
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs—0
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion—0
Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping—0
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification—0
Maximizing Confidence Alone Improves Reasoning—0
Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients—0
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving—0
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach—0
Improve Mathematical Reasoning in Language Models by Automated Process Supervision—0
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs—0
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time—0
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs—0
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference—0
Model Unlearning via Sparse Autoencoder Subspace Guided Projections—0
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization—0
Guideline Forest: Experience-Induced Multi-Guideline Reasoning with Stepwise Aggregation—0
Multi-Reference Preference Optimization for Large Language Models—0
Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision—0
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements—0
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems—0
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference—0
From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting—0
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education—0
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute—0
First-Step Advantage: Importance of Starting Right in Multi-Step Math Reasoning—0
Show:102550
← PrevPage 10 of 18Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1XolverAccuracy98.1—Unverified
2Orange-mini0-shot MRR98—Unverified