SOTAVerified

Can LLMs understand Math? -- Exploring the Pitfalls in Mathematical Reasoning

2025-05-21Unverified0· sign in to hype

Tiasa Singha Roy, Aditeya Baral, Ayush Rajesh Jhaveri, Yusuf Baig

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.

Tasks

Reproductions