SOTAVerified

Automated Theorem Proving

The goal of Automated Theorem Proving is to automatically generate a proof, given a conjecture (the target theorem) and a knowledge base of known facts, all expressed in a formal language. Automated Theorem Proving is useful in a wide range of applications, including the verification and synthesis of software and hardware systems.

Source: Learning to Prove Theorems by Learning to Generate Theorems

Papers

Showing 151–160 of 288 papers

TitleStatusHype
Planning as Theorem Proving with Heuristics—0
Probabilistic unifying relations for modelling epistemic and aleatoric uncertainty: semantics and automated reasoning with theorem proving—0
Can neural networks do arithmetic? A survey on the elementary numerical skills of state-of-the-art deep learning models—0
Lemmas: Generation, Selection, ApplicationCode0
Proceedings 11th International Workshop on Theorem Proving Components for Educational Software—0
Magnushammer: A Transformer-Based Approach to Premise Selection—0
Anti-unification and Generalization: A Survey—0
EuclidNet: Deep Visual Reasoning for Constructible Problems in Geometry—0
Solving Quantified Modal Logic Problems by Translation to Classical LogicsCode0
Keyword-based Natural Language Premise Selection for an Automatic Mathematical Statement Proving—0
Show:102550
← PrevPage 16 of 29Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Kimina-Prover-Previewcumulative80.74—Unverified
2ProofAugcumulative66—Unverified
3DeepSeek-Prover-V1.5cumulative63.5—Unverified
4Subgoal-XLcumulative56.1—Unverified
5DeepSeek-Provercumulative52—Unverified
6Lyra + GPT-4cumulative47.1—Unverified
7LEGO-Prover ChatGPTcumulative47.1—Unverified
8Decomposing the Enigmacumulative45.5—Unverified
9Evaristecumulative41—Unverified
10Evariste-7dcumulative40.6—Unverified
#ModelMetricClaimedVerifiedStatus
1EvaristePass@6458.6—Unverified
2LEGO-Prover ChatGPTPass@10057—Unverified
3Lyra + GPT-4Pass@10052—Unverified
4Evariste-7dPass@6447.5—Unverified
5GPT-fPass@6447.3—Unverified
6Evariste-1dPass@6446.7—Unverified
7DSP (62B Minerva informal)Pass@10043.9—Unverified
8Lean GPT-fPass@829.3—Unverified
9Lean tidyPass@116.8—Unverified
10Metamath GPT-fPass@82—Unverified
#ModelMetricClaimedVerifiedStatus
1MPNN-DagLSTMClassification Accuracy0.92—Unverified
2FormulaNetClassification Accuracy0.9—Unverified
3FormulaNet-basicClassification Accuracy0.89—Unverified
4Siamese 1D CNN-LSTMClassification Accuracy0.83—Unverified
5Siamese 1D CNNClassification Accuracy0.82—Unverified
#ModelMetricClaimedVerifiedStatus
14-hop GNN, sub-expression sharingPercentage correct49.95—Unverified
2Tactic Dependent LoopPercentage correct38.88—Unverified
3BoW2 (extra -ves)Percentage correct36.55—Unverified
4Deeper Wider WaveNetPercentage correct32.65—Unverified
#ModelMetricClaimedVerifiedStatus
1FormulaNetClassification Accuracy0.9—Unverified
2FormulaNet-basicClassification Accuracy0.89—Unverified
31D CNNClassification Accuracy0.83—Unverified
41D CNN-LSTMClassification Accuracy0.83—Unverified
#ModelMetricClaimedVerifiedStatus
1EvaristePass@3272.4—Unverified
2GPT-fPercentage correct56.2—Unverified
3MetaGen-IL + HolophrasmPercentage correct22.1—Unverified
4HolophrasmPercentage correct14.3—Unverified
#ModelMetricClaimedVerifiedStatus
1Evariste-7dPass@6442.5—Unverified
2Evariste-1dPass@6433.6—Unverified
3EvaristePass@6432.1—Unverified
4GPT-fPass@6430.6—Unverified
#ModelMetricClaimedVerifiedStatus
1Proverbot9001Percentage correct19.36—Unverified
2CoqGym/ASTacticPercentage correct4.99—Unverified
#ModelMetricClaimedVerifiedStatus
1ASTacticPercentage correct12.2—Unverified