SOTAVerified

Benchmarking

Papers

Showing 51515160 of 5548 papers

TitleStatusHype
On Using Distribution-Based Compositionality Assessment to Evaluate Compositional Generalisation in Machine TranslationCode0
Are Large Language Models Good at Utility Judgments?Code0
Benchmarking Language-agnostic Intent Classification for Virtual Assistant PlatformsCode0
Distributed Non-Convex Optimization with Sublinear Speedup under Intermittent Client AvailabilityCode0
VitaGraph: Building a Knowledge Graph for Biologically Relevant Learning TasksCode0
Dissecting Sample Hardness: A Fine-Grained Analysis of Hardness Characterization Methods for Data-Centric AICode0
Dissecting Dissonance: Benchmarking Large Multimodal Models Against Self-Contradictory InstructionsCode0
DispBench: Benchmarking Disparity Estimation to Synthetic CorruptionsCode0
OpenBioLink: A benchmarking framework for large-scale biomedical link predictionCode0
DispaRisk: Auditing Fairness Through Usable InformationCode0
Show:102550
← PrevPage 516 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified