SOTAVerified

Legal Reasoning

Papers

Showing 51–75 of 92 papers

TitleStatusHype
Modelling Value-oriented Legal Reasoning in LogiKEy—0
Engineering the Law-Machine Learning Translation Problem: Developing Legally Aligned Models—0
Enhancing Logical Reasoning in Large Language Models to Facilitate Legal Applications—0
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond—0
Explainable machine learning multi-label classification of Spanish legal judgements—0
Exploiting Domain-Specific Knowledge for Judgment Prediction Is No Panacea—0
Exploring the psychology of LLMs' Moral and Legal Reasoning—0
Formalising Anti-Discrimination Law in Automated Decision Systems—0
IndianBailJudgments-1200: A Multi-Attribute Dataset for Legal NLP on Indian Bail Orders—0
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding—0
KRAG Framework for Enhancing LLMs in the Legal Domain—0
LAPIS: Language Model-Augmented Police Investigation System—0
LAR-ECHR: A New Legal Argument Reasoning Task and Dataset for Cases of the European Court of Human Rights—0
Large Language Models Acing Chartered Accountancy—0
Large Language Models in Cryptocurrency Securities Cases: Can a GPT Model Meaningfully Assist Lawyers?—0
Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans—0
Law to Binary Tree -- An Formal Interpretation of Legal Natural Language—0
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models—0
LegalBench: Prototyping a Collaborative Benchmark for Legal Reasoning—0
An Argumentation-Based Legal Reasoning Approach for DL-Ontology—0
Legal Evalutions and Challenges of Large Language Models—0
Designing Normative Theories for Ethical and Legal Reasoning: LogiKEy Framework, Methodology, and Tool SupportCode0
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language ModelsCode0
Claim Extraction and Law Matching for COVID-19-related LegislationCode0
Passing the Brazilian OAB Exam: data preparation and some experimentsCode0
Show:102550
← PrevPage 3 of 4Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4Balanced Accuracy82.9—Unverified
2GPT-3.5Balanced Accuracy60.9—Unverified
3Claude-1Balanced Accuracy58.1—Unverified
#ModelMetricClaimedVerifiedStatus
1GPT-4Balanced Accuracy59.2—Unverified