SOTAVerified

The Open Verification Layer for ML Research

Community benchmark tracking and reproducibility verification. Built for researchers and autonomous research agents.

510,095 papers251,776 code links4,818 tasks

Papers

Showing 1000110025 of 510095 papers

TitleStatusHype
When Do We Need LLMs? A Diagnostic for Language-Driven Bandits0
The Missing Knowledge Layer in Cognitive Architectures for AI Agents0
Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework0
iTRIALSPACE: Programmable Virtual Lesion Trials for Controlled Evaluation of Lung CT Models0
No One Knows the State of the Art in Geospatial Foundation Models0
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools0
Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build0
Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency0
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization0
A Definition of Good Explanations and the Challenges Explaining LLM Outputs0
Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions0
Improved Knowledge Distillation for Land-Use Image Classification0
Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method0
Relational Structural Causal Models0
Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents0
From Simulation to the Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting0
Identification and Inference for Algorithmic Frontiers with Selective Labels0
Explainable Task-Oriented Token Communication for AI-Native 6G Networks0
S23DR 2026: End-to-End 3D Wireframe Prediction via DETR-Style Set Prediction with Contrastive Denoising0
JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics0
A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development0
Combining Retrieval-Augmented Text Generation with LLMs for Reading Content Recommendations0
Quantum Machine Learning for Industrial Applications0
Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation0
Leptomeningeal Collateral Detection on DSA via Vessel-Graph Neural Networks0
Show:102550
← PrevPage 401 of 20404Next →