SOTAVerified

Benchmarking

Papers

Showing 651675 of 5548 papers

TitleStatusHype
Global Rice Multi-Class Segmentation Dataset (RiceSEG): A Comprehensive and Diverse High-Resolution RGB-Annotated Images for the Development and Benchmarking of Rice Segmentation Algorithms0
When Reasoning Meets Compression: Benchmarking Compressed Large Reasoning Models on Complex Reasoning Tasks0
Benchmarking the Spatial Robustness of DNNs via Natural and Adversarial Localized Corruptions0
FIORD: A Fisheye Indoor-Outdoor Dataset with LIDAR Ground Truth for 3D Scene Reconstruction and Benchmarking0
BlenderGym: Benchmarking Foundational Model Systems for Graphics EditingCode1
Horizon Scans can be accelerated using novel information retrieval and artificial intelligence tools0
Accelerating IoV Intrusion Detection: Benchmarking GPU-Accelerated vs CPU-Based ML Libraries0
Benchmarking Synthetic Tabular Data: A Multi-Dimensional Evaluation FrameworkCode2
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models0
TDBench: Benchmarking Vision-Language Models in Understanding Top-Down ImagesCode0
Scaling Up Resonate-and-Fire Networks for Fast Deep LearningCode0
Benchmarking Federated Machine Unlearning methods for Tabular Data0
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models0
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Metacognitive Cultural Intelligence with CQ-BenchCode0
LOCO-EPI: Leave-one-chromosome-out (LOCO) as a benchmarking paradigm for deep learning based prediction of enhancer-promoter interactionsCode0
On Benchmarking Code LLMs for Android Malware Analysis0
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research PapersCode1
Towards Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-critical Scenarios0
Uni-Render: A Unified Accelerator for Real-Time Rendering Across Diverse Neural Renderers0
Simple Feedfoward Neural Networks are Almost All You Need for Time Series Forecasting0
Benchmarking Systematic Relational Reasoning with Large Language and Reasoning Models0
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation0
Unsupervised Anomaly Detection in Multivariate Time Series across Heterogeneous DomainsCode0
RL2Grid: Benchmarking Reinforcement Learning in Power Grid Operations0
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis0
Show:102550
← PrevPage 27 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified