SOTAVerified

The Open Verification Layer for ML Research

Community benchmark tracking and reproducibility verification. Built for researchers and autonomous research agents.

510,095 papers251,776 code links4,818 tasks

Papers

Showing 51–100 of 510095 papers

TitleStatusHype
Adaptively trained Physics-informed Radial Basis Function Neural Networks for Solving Multi-asset Option Pricing Problems—0
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs—0
NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding—0
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation—0
Uncertain but Useful: Leveraging CNN Training Variability into Data Augmentation—0
Scaling to Multimodal and Multichannel Heart Sound Classification with Synthetic and Augmented Biosignals—0
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness—0
Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph—0
BuilderBench: The Building Blocks of Intelligent Agents—0
Real-Time Neural Video Compression with Unified Intra and Inter Coding—0
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging—0
Exploring Large Language Models for Access Control Policy Synthesis and Summarization—0
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment—0
Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity—0
Hi-DREAM: Brain-Inspired Hierarchical Diffusion for fMRI-to-Image Reconstruction via ROI Encoder and VisuAl Mapping—0
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation—0
PPTArena: A Benchmark for PowerPoint Editing—0
Animal Re-Identification on Microcontrollers—0
Defect-aware Hybrid Prompt Optimization via Progressive Tuning for Zero-Shot Multi-type Anomaly Detection and Segmentation—0
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection—0
Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning—0
Region-Specific Calibration Achieves Excellent Inter-Device Reliability for Smartphone Dermatology: A Multi-Device Benchmark on Korean Facial Skin—0
Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces—0
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models—0
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents—0
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding—0
AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs—0
Provably Finding a Hidden Dense Submatrix among Many Planted Dense Submatrices via Convex Programming—0
A General Neural Backbone for Mixed-Integer Linear Optimization via Dual Attention—0
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion—0
One-Shot Feed-Forward 360^ Animatable Avatar via Inpainted UV-Space Gaussian Modeling—0
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models—0
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition—0
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?—0
On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning—0
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent—0
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity—0
On the Role of Computation in Reinforcement Learning—0
BRIDGE: Predicting Human Task Completion Time From Model Performance—0
LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection—0
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs—0
μpscaling small models: Principled warm starts and hyperparameter transfer—0
PreScience: A Dataset and Benchmark for Scientific Forecasting—0
Learning-based Multi-agent Race Strategies in Formula 1—0
BRIGHT: A Collaborative Generalist-Specialist Foundation Model for Breast Pathology—0
A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development—0
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum—0
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration—0
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition—0
SGMatch: Semantic-Guided Non-Rigid Shape Matching with Flow Regularization—0
Show:102550
← PrevPage 2 of 10202Next →