SOTAVerified

Benchmarking

Papers

Showing 52615270 of 5548 papers

TitleStatusHype
Conformal Prediction: A Theoretical Note and Benchmarking Transductive Node Classification in GraphsCode0
Agentic-HLS: An agentic reasoning based high-level synthesis system using large language models (AI for EDA workshop 2024)Code0
Towards Objectively Benchmarking Social Intelligence for Language Agents at Action LevelCode0
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QACode0
Custom Dual Transportation Mode Detection by Smartphone Devices Exploiting Sensor DiversityCode0
CuRe: Cultural Gaps in the Long Tail of Text-to-Image SystemsCode0
PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language ModelsCode0
CURATe: Benchmarking Personalised Alignment of Conversational AI AssistantsCode0
CUDA-GHR: Controllable Unsupervised Domain Adaptation for Gaze and Head RedirectionCode0
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise LevelsCode0
Show:102550
← PrevPage 527 of 555Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified