SOTAVerified

Benchmarking

Papers

Showing 51715180 of 5548 papers

TitleStatusHype
DFEE: Interactive DataFlow Execution and Evaluation KitCode0
Towards causal benchmarking of bias in face analysis algorithmsCode0
SORCE: Small Object Retrieval in Complex EnvironmentsCode0
Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological UnderpinningsCode0
Recognizing Object Affordances to Support Scene Reasoning for Manipulation TasksCode0
CleanPatrick: A Benchmark for Image Data CleaningCode0
Detecting critical treatment effect bias in small subgroupsCode0
AI-generated Image Quality Assessment in Visual CommunicationCode0
SOSD: A Benchmark for Learned IndexesCode0
OpenML Benchmarking SuitesCode0
Show:102550
← PrevPage 518 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified