SOTAVerified

Benchmarking

Papers

Showing 18011825 of 5548 papers

TitleStatusHype
LEXam: Benchmarking Legal Reasoning on 340 Law Exams0
Benchmarking MOEAs for solving continuous multi-objective RL problemsCode0
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference0
CompBench: Benchmarking Complex Instruction-guided Image Editing0
ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models0
Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind0
Disambiguation in Conversational Question Answering in the Era of LLM: A Survey0
OSS-Bench: Benchmark Generator for Coding LLMsCode0
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation0
Machine Learning-Based Analysis of ECG and PCG Signals for Rheumatic Heart Disease Detection: A Scoping Review (2015-2025)0
GenderBench: Evaluation Suite for Gender Biases in LLMsCode0
SoftPQ: Robust Instance Segmentation Evaluation via Soft Matching and Tunable ThresholdsCode0
HumaniBench: A Human-Centric Framework for Large Multimodal Models EvaluationCode0
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems0
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents0
Benchmarking CFAR and CNN-based Peak Detection Algorithms in ISAC under Hardware Impairments0
Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models0
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese0
VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning TasksCode0
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale0
STEP: A Unified Spiking Transformer Evaluation Platform for Fair and Reproducible BenchmarkingCode0
CleanPatrick: A Benchmark for Image Data CleaningCode0
Visual Anomaly Detection under Complex View-Illumination Interplay: A Large-Scale Benchmark0
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models0
Relation Extraction Across Entire Books to Reconstruct Community Networks: The AffilKG Datasets0
Show:102550
← PrevPage 73 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified