SOTAVerified

Sequential Decision Making

Papers

Showing 51100 of 1210 papers

TitleStatusHype
Deep Reinforcement Learning based Recommendation with Explicit User-Item Interactions ModelingCode1
Comparing Deep Reinforcement Learning Algorithms in Two-Echelon Supply ChainsCode1
Deep Reinforcement Learning for Entity AlignmentCode1
Can language agents be alternatives to PPO? A Preliminary Empirical Study On OpenAI GymCode1
Learning Dynamic Belief Graphs to Generalize on Text-Based GamesCode1
Large Language Models for Planning: A Comprehensive and Systematic SurveyCode1
The Sandbox Environment for Generalizable Agent Research (SEGAR)Code1
Thinking Fast and Slow with Deep Learning and Tree SearchCode1
Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy OptimizationCode1
Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingCode1
Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision MakingCode1
Unified Models of Human Behavioral Agents in Bandits, Contextual Bandits and RLCode1
Layered and Staged Monte Carlo Tree Search for SMT Strategy SynthesisCode1
Learning Multi-Level Hierarchies with HindsightCode1
Free from Bellman Completeness: Trajectory Stitching via Model-based Return-conditioned Supervised LearningCode1
Hybrid Multi-agent Deep Reinforcement Learning for Autonomous Mobility on Demand SystemsCode1
Enabling Intelligent Interactions between an Agent and an LLM: A Reinforcement Learning ApproachCode1
Extracting Reward Functions from Diffusion ModelsCode1
Independent Reinforcement Learning for Weakly Cooperative Multiagent Traffic Control ProblemCode1
How Can LLM Guide RL? A Value-Based ApproachCode1
Effective Reinforcement Learning through Evolutionary Surrogate-Assisted PrescriptionCode1
Approximate Inference in Discrete Distributions with Monte Carlo Tree Search and Value FunctionsCode1
Breadcrumbs to the Goal: Goal-Conditioned Exploration from Human-in-the-Loop FeedbackCode1
Large Language Model as a Policy Teacher for Training Reinforcement Learning AgentsCode1
Efficient Nonmyopic Bayesian Optimization via One-Shot Multi-Step TreesCode1
AdaPlanner: Adaptive Planning from Feedback with Language ModelsCode1
Can Increasing Input Dimensionality Improve Deep Reinforcement Learning?Code1
Bridging POMDPs and Bayesian decision making for robust maintenance planning under model uncertainty: An application to railway systemsCode1
ContainerGym: A Real-World Reinforcement Learning Benchmark for Resource AllocationCode1
CertRL: Formalizing Convergence Proofs for Value and Policy Iteration in CoqCode1
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning PoliciesCode1
Dynamic Causal Bayesian OptimizationCode1
RELIEF: Reinforcement Learning Empowered Graph Feature Prompt TuningCode1
Dynamic Multi-Robot Task Allocation under Uncertainty and Temporal ConstraintsCode1
Efficient Symptom Inquiring and Diagnosis via Adaptive Alignment of Reinforcement Learning and ClassificationCode1
An Alternative Softmax Operator for Reinforcement LearningCode1
IQ-Learn: Inverse soft-Q Learning for ImitationCode1
Object-Aware Regularization for Addressing Causal Confusion in Imitation LearningCode1
On Generalization Across Environments In Multi-Objective Reinforcement LearningCode1
Out of the Cage: How Stochastic Parrots Win in Cyber Security EnvironmentsCode1
Counterfactual Explanations in Sequential Decision Making Under UncertaintyCode1
LLF-Bench: Benchmark for Interactive Learning from Language FeedbackCode1
An empirical evaluation of active inference in multi-armed banditsCode1
Curriculum-based Reinforcement Learning for Distribution System Critical Load RestorationCode1
Reinforcement Learning for Temporal Logic Control Synthesis with Probabilistic Satisfaction GuaranteesCode1
Decision Stacks: Flexible Reinforcement Learning via Modular Generative ModelsCode1
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State SpacesCode1
Reinforcement learning with combinatorial actions for coupled restless banditsCode1
RLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement LearningCode1
Distance Weighted Supervised Learning for Offline Interaction DataCode0
Show:102550
← PrevPage 2 of 25Next →

No leaderboard results yet.