SOTAVerified

Benchmarking

Papers

Showing 52915300 of 5548 papers

TitleStatusHype
Architecture Analysis and Benchmarking of 3D U-shaped Deep Learning Models for Thoracic Anatomical SegmentationCode0
XCompress: LLM assisted Python-based text compression toolkitCode0
A Framework for Generating Informative Benchmark InstancesCode0
What's Different between Visual Question Answering for Machine "Understanding" Versus for Accessibility?Code0
Towards Robust Metrics for Concept Representation EvaluationCode0
Statistical Multicriteria Evaluation of LLM-Generated TextCode0
ANTHROPOS-V: benchmarking the novel task of Crowd Volume EstimationCode0
Answer Consolidation: Formulation and BenchmarkingCode0
A Benchmark on Extremely Weakly Supervised Text Classification: Reconcile Seed Matching and Prompting ApproachesCode0
A novel evaluation methodology for supervised Feature Ranking algorithmsCode0
Show:102550
← PrevPage 530 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified