SOTAVerified

Benchmarking

Papers

Showing 50915100 of 5548 papers

TitleStatusHype
Single and Multi-Hop Question-Answering Datasets for Reticular Chemistry with GPT-4-TurboCode0
Benchmarking machine learning for bowel sound pattern classification from tabular features to pretrained modelsCode0
On dataset transferability in medical image classificationCode0
Are Synthetic Corruptions A Reliable Proxy For Real-World Corruptions?Code0
Do LLM Evaluators Prefer Themselves for a Reason?Code0
YOLOBench: Benchmarking Efficient Object Detectors on Embedded SystemsCode0
Benchmarking Long-tail Generalization with Likelihood SplitsCode0
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and BenchmarkingCode0
On Empirical Comparisons of Optimizers for Deep LearningCode0
Benchmarking LLMs' Judgments with No Gold StandardCode0
Show:102550
← PrevPage 510 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified