SOTAVerified

Benchmarking

Papers

Showing 626650 of 5548 papers

TitleStatusHype
COCO: The Large Scale Black-Box Optimization Benchmarking (bbob-largescale) Test SuiteCode1
Coarse-to-Fine Q-attention with Learned Path RankingCode1
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial OptimizationCode1
Codabench: Flexible, Easy-to-Use and Reproducible Benchmarking PlatformCode1
Clinical Prompt Learning with Frozen Language ModelsCode1
CLoG: Benchmarking Continual Learning of Image Generation ModelsCode1
ClearPose: Large-scale Transparent Object Dataset and BenchmarkCode1
ClimART: A Benchmark Dataset for Emulating Atmospheric Radiative Transfer in Weather and Climate ModelsCode1
CloudEval-YAML: A Practical Benchmark for Cloud Configuration GenerationCode1
CODEBench: A Neural Architecture and Hardware Accelerator Co-Design FrameworkCode1
Collective Knowledge: organizing research projects as a database of reusable components and portable workflows with common APIsCode1
CovDocker: Benchmarking Covalent Drug Design with Tasks, Datasets, and SolutionsCode1
AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite ImageryCode1
CIBench: Evaluating Your LLMs with a Code Interpreter PluginCode1
CIDEr: Consensus-based Image Description EvaluationCode1
CheXphoto: 10,000+ Photos and Transformations of Chest X-rays for Benchmarking Deep Learning RobustnessCode1
CHILI: Chemically-Informed Large-scale Inorganic Nanomaterials Dataset for Advancing Graph Machine LearningCode1
CIPCaD-Bench: Continuous Industrial Process datasets for benchmarking Causal Discovery methodsCode1
Align and Distill: Unifying and Improving Domain Adaptive Object DetectionCode1
On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic WritingCode1
BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language modelsCode1
CheX-GPT: Harnessing Large Language Models for Enhanced Chest X-ray Report LabelingCode1
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learningCode1
CCTV-Gun: Benchmarking Handgun Detection in CCTV ImagesCode1
Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive?Code1
Show:102550
← PrevPage 26 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified