SOTAVerified

Automated Theorem Proving

The goal of Automated Theorem Proving is to automatically generate a proof, given a conjecture (the target theorem) and a knowledge base of known facts, all expressed in a formal language. Automated Theorem Proving is useful in a wide range of applications, including the verification and synthesis of software and hardware systems.

Source: Learning to Prove Theorems by Learning to Generate Theorems

Papers

Showing 1–50 of 288 papers

TitleStatusHype
CriticLean: Critic-Guided Reinforcement Learning for Mathematical FormalizationCode1
Prover Agent: An Agent-based Framework for Formal Mathematical Proofs—0
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving—0
MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?—0
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal VerificationCode1
LeanExplore: A search engine for Lean 4 declarationsCode2
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening—0
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations—0
ProofNet++: A Neuro-Symbolic System for Formal Proof Verification with Self-Correction—0
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement LearningCode1
Autoformalization in the Era of Large Language Models: A SurveyCode5
RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation—0
Enumerate-Conjecture-Prove: Formally Solving Answer-Construction Problems in Math CompetitionsCode0
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement—0
MIRB: Mathematical Information Retrieval BenchmarkCode0
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on InequalitiesCode1
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data GenerationCode0
MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation—0
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning—0
Proceedings The 13th International Workshop on Theorem proving components for Educational software—0
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving—0
DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal DecompositionCode5
The Limits of AI Explainability: An Algorithmic Information Theory Approach—0
APE-Bench I: Towards File-level Automated Proof Engineering of Formal Math Libraries—0
Hua-Chen New Theory of Economic Optimization—0
Hierarchical Attention Generates Better ProofsCode0
Neural Theorem Proving: Generating and Structuring Proofs for Formal Verification—0
Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement LearningCode3
Reasoning Models Can Be Effective Without Thinking—0
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection—0
Leanabell-Prover: Posttraining Scaling in Formal ReasoningCode1
Reasoning Under Threat: Symbolic and Neural Techniques for Cybersecurity Verification—0
A Survey on Mathematical Reasoning and Optimization with Large Language ModelsCode0
Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview—0
Local Look-Ahead Guidance via Verifier-in-the-Loop for Automated Theorem Proving—0
Efficient Neural Clause-Selection Reinforcement—0
MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem ProvingCode1
Faithful Logic Embeddings in HOL -- Deep and Shallow—0
Quantum Machine Learning in Precision Medicine and Drug Discovery -- A Game Changer for Tailored Treatments?—0
LeanProgress: Guiding Search for Neural Theorem Proving via Proof Progress Prediction—0
A Combinatorial Identities Benchmark for Theorem Proving via Automated Theorem Generation—0
Activation Steering in Neural Theorem Provers—0
Generating Millions Of Lean Theorems With Proofs By Exploring State Transition Graphs—0
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem ProvingCode3
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation—0
ProofWala: Multilingual Proof Data Synthesis and Theorem-ProvingCode1
BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving—0
STP: Self-play LLM Theorem Provers with Iterative Conjecturing and ProvingCode2
Efficient Neural Theorem Proving via Fine-grained Proof Structure AnalysisCode1
LemmaHead: RAG Assisted Proof Generation Using Large Language Models—0
Show:102550
← PrevPage 1 of 6Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Kimina-Prover-Previewcumulative80.74—Unverified
2ProofAugcumulative66—Unverified
3DeepSeek-Prover-V1.5cumulative63.5—Unverified
4Subgoal-XLcumulative56.1—Unverified
5DeepSeek-Provercumulative52—Unverified
6Lyra + GPT-4cumulative47.1—Unverified
7LEGO-Prover ChatGPTcumulative47.1—Unverified
8Decomposing the Enigmacumulative45.5—Unverified
9Evaristecumulative41—Unverified
10Evariste-7dcumulative40.6—Unverified
#ModelMetricClaimedVerifiedStatus
1EvaristePass@6458.6—Unverified
2LEGO-Prover ChatGPTPass@10057—Unverified
3Lyra + GPT-4Pass@10052—Unverified
4Evariste-7dPass@6447.5—Unverified
5GPT-fPass@6447.3—Unverified
6Evariste-1dPass@6446.7—Unverified
7DSP (62B Minerva informal)Pass@10043.9—Unverified
8Lean GPT-fPass@829.3—Unverified
9Lean tidyPass@116.8—Unverified
10Metamath GPT-fPass@82—Unverified
#ModelMetricClaimedVerifiedStatus
1MPNN-DagLSTMClassification Accuracy0.92—Unverified
2FormulaNetClassification Accuracy0.9—Unverified
3FormulaNet-basicClassification Accuracy0.89—Unverified
4Siamese 1D CNN-LSTMClassification Accuracy0.83—Unverified
5Siamese 1D CNNClassification Accuracy0.82—Unverified
#ModelMetricClaimedVerifiedStatus
14-hop GNN, sub-expression sharingPercentage correct49.95—Unverified
2Tactic Dependent LoopPercentage correct38.88—Unverified
3BoW2 (extra -ves)Percentage correct36.55—Unverified
4Deeper Wider WaveNetPercentage correct32.65—Unverified
#ModelMetricClaimedVerifiedStatus
1FormulaNetClassification Accuracy0.9—Unverified
2FormulaNet-basicClassification Accuracy0.89—Unverified
31D CNNClassification Accuracy0.83—Unverified
41D CNN-LSTMClassification Accuracy0.83—Unverified
#ModelMetricClaimedVerifiedStatus
1EvaristePass@3272.4—Unverified
2GPT-fPercentage correct56.2—Unverified
3MetaGen-IL + HolophrasmPercentage correct22.1—Unverified
4HolophrasmPercentage correct14.3—Unverified
#ModelMetricClaimedVerifiedStatus
1Evariste-7dPass@6442.5—Unverified
2Evariste-1dPass@6433.6—Unverified
3EvaristePass@6432.1—Unverified
4GPT-fPass@6430.6—Unverified
#ModelMetricClaimedVerifiedStatus
1Proverbot9001Percentage correct19.36—Unverified
2CoqGym/ASTacticPercentage correct4.99—Unverified
#ModelMetricClaimedVerifiedStatus
1ASTacticPercentage correct12.2—Unverified