SOTAVerified

Sequential Decision Making

Papers

Showing 1–25 of 1210 papers

TitleStatusHype
AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air—0
LLM-Stackelberg Games: Conjectural Reasoning Equilibria and Their Applications to Spearphishing—0
A Survey of Continual Reinforcement Learning—0
Flow-Based Single-Step Completion for Efficient and Expressive Policy Learning—0
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes—0
Efficient Strategy Synthesis for MDPs via Hierarchical Block Decomposition—0
UProp: Investigating the Uncertainty Propagation of LLMs in Multi-Step Agentic Decision-MakingCode0
Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards—0
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic EnvironmentsCode0
Common Benchmarks Undervalue the Generalization Power of Programmatic PoliciesCode0
Leveraging In-Context Learning for Language Model Agents—0
Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity—0
Towards Responsible AI: Advances in Safety, Fairness, and Accountability of Autonomous Systems—0
TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement LearningCode0
How to Provably Improve Return Conditioned Supervised Learning?—0
QForce-RL: Quantized FPGA-Optimized Reinforcement Learning Compute Engine—0
Contextual Experience Replay for Self-Improvement of Language Agents—0
AutoQD: Automatic Discovery of Diverse Behaviors with Quality-Diversity Optimization—0
TextAtari: 100K Frames Game Playing with Language AgentsCode0
Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation—0
Emergent Risk Awareness in Rational Agents under Resource Constraints—0
Adaptive Frontier Exploration on Graphs with Applications to Network-Based Disease Testing—0
Variational Deep Learning via Implicit Regularization—0
Large Language Models for Planning: A Comprehensive and Systematic SurveyCode1
DDO: Dual-Decision Optimization via Multi-Agent Collaboration for LLM-Based Medical Consultation—0
Show:102550
← PrevPage 1 of 49Next →

No leaderboard results yet.