SOTAVerified

Benchmarking

Papers

Showing 51215130 of 5548 papers

TitleStatusHype
Towards a Comprehensive Benchmark for Pathological Lymph Node Metastasis in Breast Cancer SectionsCode0
Benchmarking Large Language Model Uncertainty for Prompt OptimizationCode0
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining DatasetsCode0
On the Evaluation Consistency of Attribution-based ExplanationsCode0
On the Evaluation of Conditional GANsCode0
A Classification Benchmark for Artificial Intelligence Detection of Laryngeal Cancer from Patient VoiceCode0
Arena-Rosnav 2.0: A Development and Benchmarking Platform for Robot Navigation in Highly Dynamic EnvironmentsCode0
On the Fragility of Active Learners for Text ClassificationCode0
Distributing Deep Learning Hyperparameter Tuning for 3D Medical Image SegmentationCode0
Benchmarking Large Language Models on Communicative Medical Coaching: a Novel System and DatasetCode0
Show:102550
← PrevPage 513 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified