SOTAVerified

Benchmarking

Papers

Showing 19011925 of 5548 papers

TitleStatusHype
Benchmarking GNNs Using Lightning Network Data0
Benchmarking structure-based three-dimensional molecular generative models using GenBench3D: ligand conformation quality mattersCode1
From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano0
Towards Stable 3D Object Detection0
SH17: A Dataset for Human Safety and Personal Protective Equipment Detection in Manufacturing IndustryCode2
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation0
Craftium: An Extensible Framework for Creating Reinforcement Learning EnvironmentsCode2
Benchmarking Complex Instruction-Following with Multiple Constraints CompositionCode2
Benchmark on Drug Target Interaction Modeling from a Structure PerspectiveCode1
Benchmarking End-To-End Performance of AI-Based Chip Placement Algorithms0
Comics Datasets Framework: Mix of Comics datasets for detection benchmarkingCode1
Social Bias in Large Language Models For Bangla: An Empirical Study on Gender and Religious BiasCode0
CoIR: A Comprehensive Benchmark for Code Information Retrieval ModelsCode2
GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language ModelsCode1
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking DatasetCode1
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations0
Evaluating the Ability of LLMs to Solve Semantics-Aware Process Mining TasksCode0
Open foundation models for Azerbaijani language0
Occlusion-Aware Seamless SegmentationCode1
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile AgentsCode1
Modified CMA-ES Algorithm for Multi-Modal Optimization: Incorporating Niching Strategies and Dynamic Adaptation Mechanism0
Task-oriented Over-the-air Computation for Edge-device Co-inference with Balanced Classification Accuracy0
MIRAI: Evaluating LLM Agents for Event Forecasting0
BERGEN: A Benchmarking Library for Retrieval-Augmented GenerationCode3
ProductAgent: Benchmarking Conversational Product Search Agent with Asking Clarification Questions0
Show:102550
← PrevPage 77 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified