SOTAVerified

Benchmarking

Papers

Showing 53015310 of 5548 papers

TitleStatusHype
Benchmarking Foundation Models on Exceptional Cases: Dataset Creation and ValidationCode0
CSS: A Large-scale Cross-schema Chinese Text-to-SQL Medical DatasetCode0
Cryo-RALib -- a modular library for accelerating alignment in cryo-EMCode0
What the Weight?! A Unified Framework for Zero-Shot Knowledge CompositionCode0
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive ProgressionsCode0
Cross-Lingual Text Classification of Transliterated Hindi and MalayalamCode0
Benchmarking Flexible Electric Loads Scheduling Algorithms under Market Price UncertaintyCode0
Yum-me: A Personalized Nutrient-based Meal Recommender SystemCode0
Benchmarking Federated Learning for Semantic Datasets: Federated Scene Graph GenerationCode0
Cross-lingual sentiment classification in low-resource Bengali languageCode0
Show:102550
← PrevPage 531 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified