SOTAVerified

Benchmarking

Papers

Showing 51015110 of 5548 papers

TitleStatusHype
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and ReasoningCode0
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence BenchmarksCode0
Benchmarking LLM-based Relevance Judgment MethodsCode0
Toward 3D Object Reconstruction from Stereo ImagesCode0
DLAMA: A Framework for Curating Culturally Diverse Facts for Probing the Knowledge of Pretrained Language ModelsCode0
Skelite: Compact Neural Networks for Efficient Iterative SkeletonizationCode0
Divergent Creativity in Humans and Large Language ModelsCode0
A Kernel-Based Approach for Accurate Steady-State Detection in Performance Time SeriesCode0
A Closer Look at Temporal Sentence Grounding in Videos: Dataset and MetricCode0
Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue SystemsCode0
Show:102550
← PrevPage 511 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified