SOTAVerified

Math Word Problem Solving

A math word problem is a mathematical exercise (such as in a textbook, worksheet, or exam) where significant background information on the problem is presented in ordinary language rather than in mathematical notation. As most word problems involve a narrative of some sort, they are sometimes referred to as story problems and may vary in the amount of technical language used.

Papers

Showing 41–50 of 107 papers

TitleStatusHype
FinanceMath: Knowledge-Intensive Math Reasoning in Finance DomainsCode1
MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem SolvingCode1
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language ReasoningCode1
Automatic Model Selection with Large Language Models for ReasoningCode1
MWPToolkit: An Open-Source Framework for Deep Learning-Based Math Word Problem SolversCode1
Graph-to-Tree Neural Networks for Learning Structured Input-Output Translation with Applications to Semantic Parsing and Math Word ProblemCode1
Graph-to-Tree Learning for Solving Math Word ProblemsCode1
Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsCode1
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word ProblemsCode1
Augmenting Math Word Problems via Iterative Question ComposingCode1
Show:102550
← PrevPage 5 of 11Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Gemini 2.0 Flash ExperimentalAccuracy89.7—Unverified
2Qwen2.5-Math-72B-Instruct(TIR,Greedy)Accuracy88.1—Unverified
3GPT-4 Turbo (MACM, w/code, voting)Accuracy87.92—Unverified
4Qwen2.5-Math-72B-Instruct(COT,Greedy)Accuracy85.9—Unverified
5Qwen2.5-Math-7B-Instruct(TIR,Greedy)Accuracy85.2—Unverified
6GPT-4-code model (CSV, w/ code, SC, k=16)Accuracy84.3—Unverified
7Qwen2-Math-72B-Instruct(greedy)Accuracy84—Unverified
8Qwen2.5-Math-7B-Instruct(COT,Greedy)Accuracy83.6—Unverified
9Qwen2.5-Math-1.5B-Instruct(TIR,Greedy)Accuracy79.9—Unverified
10OpenMath2-Llama3.1-70B (majority@256)Accuracy79.6—Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4 DUPAccuracy94.2—Unverified
2GPT-4 (Teaching-Inspired)Execution Accuracy93.9—Unverified
3GPT-4 (Model Selection)Execution Accuracy93.7—Unverified
4Qwen2(CoT + Code Interpreter)Execution Accuracy92.3—Unverified
5GPT-4 (PHP)Execution Accuracy91.9—Unverified
6OpenMath-CodeLlama-70B (w/ code)Execution Accuracy87.8—Unverified
7MathCoder-L-70BExecution Accuracy84.9—Unverified
8PoT_Eng (self-consistency @ 5)Execution Accuracy83.7—Unverified
9CoT_Eng (self-consistency @ 5)Execution Accuracy82.5—Unverified
10MMOS-CODE-34B(0-shot)Execution Accuracy80.6—Unverified
#ModelMetricClaimedVerifiedStatus
1OpenMath-CodeLlama-70B (w/ code)Accuracy (%)95.7—Unverified
2MsAT-DeductReasonerAccuracy (%)94.3—Unverified
3ATHENA (roberta-large)Accuracy (%)93—Unverified
4Exp-TreeAccuracy (%)92.3—Unverified
5Multi-viewAccuracy (%)92.3—Unverified
6ATHENA (roberta-base)Accuracy (%)92.2—Unverified
7Roberta-DeductReasonerAccuracy (%)92—Unverified
8DeBERTa (PM + VM)Accuracy (%)91—Unverified
9EPTAccuracy (%)88.7—Unverified
10Graph2Tree with RoBERTaAccuracy (%)88.7—Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4 (Teaching-Inspired)Accuracy (5-fold)94.3—Unverified
2ATHENA (roberta-large)Accuracy (training-test)86.5—Unverified
3Multi-view* (ours)Accuracy (5-fold)85.2—Unverified
4ATHENA (roberta-base)Accuracy (training-test)84.4—Unverified
5Generate and RankAccuracy (5-fold)84.3—Unverified
6Exp-TreeAccuracy (5-fold)84.1—Unverified
7REAL2: Memory-augmented SolverAccuracy (5-fold)83.18—Unverified
8Roberta-DeductReasonerAccuracy (5-fold)83—Unverified
9MWP-BERTAccuracy (5-fold)82.4—Unverified
10Recall and LearnAccuracy (5-fold)80.8—Unverified