SOTAVerified

Benchmarking

Papers

Showing 52315240 of 5548 papers

TitleStatusHype
Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training BenchmarkCode0
Deciphering the Underserved: Benchmarking LLM OCR for Low-Resource ScriptsCode0
Towards IID representation learning and its application on biomedical dataCode0
A projected nonlinear state-space model for forecasting time series signalsCode0
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech EvaluationCode0
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word ProblemCode0
Dealing with missing data using attention and latent space regularizationCode0
DCR: Quantifying Data Contamination in LLMs EvaluationCode0
DateLogicQA: Benchmarking Temporal Biases in Large Language ModelsCode0
Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing EvaluationCode0
Show:102550
← PrevPage 524 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified