SOTAVerified

Benchmarking

Papers

Showing 676700 of 5548 papers

TitleStatusHype
SimBank: from Simulation to Solution in Prescriptive Process Monitoring0
Generalization Bias in Large Language Model Summarization of Scientific Research0
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric VideosCode1
Why Stop at One Error? Benchmarking LLMs as Data Science Code Debuggers for Multi-Hop and Multi-Bug ErrorsCode0
Benchmarking Ultra-Low-Power μNPUs0
An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model0
LIM: Large Interpolator Model for Dynamic Reconstruction0
Assessing Foundation Models for Sea Ice Type Segmentation in Sentinel-1 SAR Imagery0
Benchmarking Deep Learning-Based Methods for Irradiance Nowcasting with Sky Images0
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?Code0
Evaluating Text-to-Image Synthesis with a Conditional Fréchet Distance0
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition0
GateLens: A Reasoning-Enhanced LLM Agent for Automotive Software Release Analytics0
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsCode1
A Comprehensive Benchmark for RNA 3D Structure-Function ModelingCode1
CSPO: Cross-Market Synergistic Stock Price Movement Forecasting with Pseudo-volatility Optimization0
Can geometric combinatorics improve RNA branching predictions?Code0
RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy0
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIsCode3
Benchmarking and optimizing organism wide single-cell RNA alignment methodsCode0
TerraTorch: The Geospatial Foundation Models ToolkitCode4
Benchmarking Machine Learning Methods for Distributed Acoustic Sensing0
Reservoir Computing with a Single Oscillating Gas Bubble: Emphasizing the Chaotic Regime0
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy0
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language ModelsCode1
Show:102550
← PrevPage 28 of 222Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1GPT-4 TurboACC0.56Unverified