SOTAVerified

Offline RL

Papers

Showing 110 of 755 papers

TitleStatusHype
From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning0
Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMsCode0
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning0
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning0
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL0
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using SparsityCode0
CAWR: Corruption-Averse Advantage-Weighted Regression for Robust Policy OptimizationCode0
IntelliLung: Advancing Safe Mechanical Ventilation using Offline RL with Hybrid Actions and Clinically Aligned Rewards0
Toward Explainable Offline RL: Analyzing Representations in Intrinsically Motivated Decision Transformers0
DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under UncertaintyCode0
Show:102550
← PrevPage 1 of 76Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1KFCAverage Reward81.8Unverified
2ADMPOAverage Reward81Unverified
3Decision Transformer (DT)Average Reward73.5Unverified
#ModelMetricClaimedVerifiedStatus
1ParPID4RL Normalized Score151.4Unverified