SOTAVerified

Benchmarking

Papers

Showing 51315140 of 5548 papers

TitleStatusHype
Benchmarking Large Language Models for Math Reasoning TasksCode0
Benchmarking Large Language Models for Image Classification of Marine MammalsCode0
On the Loss of Context-awareness in General Instruction Fine-tuningCode0
HumaniBench: A Human-Centric Framework for Large Multimodal Models EvaluationCode0
SNaC: Coherence Error Detection for Narrative SummarizationCode0
SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking ServicesCode0
Using Motif Transitions for Temporal Graph GenerationCode0
Accurate Peak Detection in Multimodal Optimization via Approximated Landscape LearningCode0
Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious BiasCode0
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician ExamsCode0
Show:102550
← PrevPage 514 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified