| RL Dreams: Policy Gradient Optimization for Score Distillation based 3D Generation | Dec 8, 2023 | 3D GenerationDenoising | —Unverified | 0 | 0 |
| ROCM: RLHF on consistency models | Mar 8, 2025 | Policy Gradient Methods | —Unverified | 0 | 0 |
| Safe Reinforcement Learning via Projection on a Safe Set: How to Achieve Optimality? | Apr 2, 2020 | Policy Gradient MethodsQ-Learning | —Unverified | 0 | 0 |
| Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional Manifolds | Sep 25, 2023 | Policy Gradient MethodsReinforcement Learning (RL) | —Unverified | 0 | 0 |
| Sample Complexity of Policy Gradient Finding Second-Order Stationary Points | Dec 2, 2020 | Policy Gradient MethodsReinforcement Learning (RL) | —Unverified | 0 | 0 |
| Sample-efficient actor-critic algorithms with an etiquette for zero-sum Markov games | Sep 29, 2021 | Policy Gradient Methods | —Unverified | 0 | 0 |
| Sample-efficient Deep Reinforcement Learning for Dialog Control | Dec 18, 2016 | Deep Reinforcement LearningPolicy Gradient Methods | —Unverified | 0 | 0 |
| Sample Efficient Reinforcement Learning with REINFORCE | Oct 22, 2020 | Policy Gradient Methodsreinforcement-learning | —Unverified | 0 | 0 |
| Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL | Apr 8, 2019 | continuous-controlContinuous Control | —Unverified | 0 | 0 |
| Score-Aware Policy-Gradient Methods and Performance Guarantees using Local Lyapunov Conditions: Applications to Product-Form Stochastic Networks and Queueing Systems | Dec 5, 2023 | FormModel-based Reinforcement Learning | —Unverified | 0 | 0 |