SOTAVerified

Q-Learning

The goal of Q-learning is to learn a policy, which tells an agent what action to take under what circumstances.

( Image credit: Playing Atari with Deep Reinforcement Learning )

Papers

Showing 451–475 of 1918 papers

TitleStatusHype
Numeric Reward Machines—0
Reinforcement Learning Problem Solving with Large Language Models—0
Using Deep Q-Learning to Dynamically Toggle between Push/Pull Actions in Computational Trust Mechanisms—0
Q-learning with temporal memory to navigate turbulence—0
Age of Information Minimization using Multi-agent UAVs based on AI-Enhanced Mean Field Resource Allocation—0
AFU: Actor-Free critic Updates in off-policy RL for continuous controlCode0
Recursive Backwards Q-Learning in Deterministic Environments—0
Unified ODE Analysis of Smooth Q-Learning Algorithms—0
Continuous-time Risk-sensitive Reinforcement Learning via Quadratic Variation Penalty—0
Data-Incremental Continual Offline Reinforcement Learning—0
From r to Q^*: Your Language Model is Secretly a Q-Function—0
Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL—0
Advancing Forest Fire Prevention: Deep Reinforcement Learning for Effective Firebreak Placement—0
Prelimit Coupling and Steady-State Convergence of Constant-stepsize Nonsmooth Contractive SA—0
Traffic Signal Control and Speed Offset Coordination Using Q-Learning for Arterial Road Networks—0
Deep Reinforcement Learning Control for Disturbance Rejection in a Nonlinear Dynamic System with Parametric Uncertainty—0
Growing Q-Networks: Solving Continuous Control Tasks with Adaptive Control Resolution—0
Superior Genetic Algorithms for the Target Set Selection Problem Based on Power-Law Parameter Choices and Simple Greedy HeuristicsCode0
Data-Driven Knowledge Transfer in Batch Q^* Learning—0
Utilizing Maximum Mean Discrepancy Barycenter for Propagating the Uncertainty of Value Functions in Reinforcement Learning—0
EnCoMP: Enhanced Covert Maneuver Planning with Adaptive Threat-Aware Visibility Estimation using Offline Reinforcement Learning—0
From Two-Dimensional to Three-Dimensional Environment with Q-Learning: Modeling Autonomous Navigation with Reinforcement Learning and no LibrariesCode0
Compressed Federated Reinforcement Learning with a Generative ModelCode0
DASA: Delay-Adaptive Multi-Agent Stochastic Approximation—0
Semantic-Aware Remote Estimation of Multiple Markov Sources Under Constraints—0
Show:102550
← PrevPage 19 of 77Next →

No leaderboard results yet.