SOTAVerified

Benchmarking

Papers

Showing 18261850 of 5548 papers

TitleStatusHype
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models0
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and ChallengesCode0
TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMsCode0
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems0
Visual Fidelity Index for Generative Semantic Communications with Critical Information Embedding0
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto LanguageCode0
JointDistill: Adaptive Multi-Task Distillation for Joint Depth Estimation and Scene Segmentation0
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs0
Real-World fNIRS-Based Brain-Computer Interfaces: Benchmarking Deep Learning and Classical Models in Interactive Gaming0
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs0
Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization0
GNN-Suite: a Graph Neural Network Benchmarking Framework for Biomedical InformaticsCode0
On the Evaluation of Engineering Artificial General Intelligence0
Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1MCode0
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models0
VeriFact: Enhancing Long-Form Factuality Evaluation with Refined Fact Extraction and Reference Facts0
RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo0
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning0
BioVFM-21M: Benchmarking and Scaling Self-Supervised Vision Foundation Models for Biomedical Image AnalysisCode0
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation0
TARGET: Benchmarking Table Retrieval for Generative Tasks0
A Standardized Benchmark Set of Clustering Problem Instances for Comparing Black-Box Optimizers0
How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference0
Grounding Synthetic Data Evaluations of Language Models in Unsupervised Document CorporaCode0
ExEBench: Benchmarking Foundation Models on Extreme Earth EventsCode0
Show:102550
← PrevPage 74 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified