SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 526–550 of 655 papers

TitleStatusHype
Remote Contextual Bandits—0
Residual Bootstrap Exploration for Bandit Algorithms—0
Revised Progressive-Hedging-Algorithm Based Two-layer Solution Scheme for Bayesian Reinforcement Learning—0
Reward Biased Maximum Likelihood Estimation for Reinforcement Learning—0
Risk and optimal policies in bandit experiments—0
Risk-averse Contextual Multi-armed Bandit Problem with Linear Payoffs—0
Risk-Constrained Thompson Sampling for CVaR Bandits—0
Robust Dynamic Assortment Optimization in the Presence of Outlier Customers—0
Robust Policy Switching for Antifragile Reinforcement Learning for UAV Deconfliction in Adversarial Environments—0
Robust Thompson Sampling Algorithms Against Reward Poisoning Attacks—0
Safe Linear Leveling Bandits—0
Safe Linear Thompson Sampling with Side Information—0
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit—0
The Price of Incentivizing Exploration: A Characterization via Thompson Sampling and Sample Complexity—0
Sampling Acquisition Functions for Batch Bayesian Optimization—0
Satisficing in Time-Sensitive Bandit Learning—0
Scalable and Interpretable Contextual Bandits: A Literature Review and Retail Offer Prototype—0
Scalable Generalized Linear Bandits: Online Computation and Hashing—0
Scalable Neural Contextual Bandit for Recommender Systems—0
Scalable regret for learning to control network-coupled subsystems with unknown dynamics—0
Scalable Thompson Sampling using Sparse Gaussian Process Models—0
Scalable Thompson Sampling via Optimal Transport—0
Scaling Multi-Armed Bandit Algorithms—0
Screening for an Infectious Disease as a Problem in Stochastic Control—0
Semi-Parametric Contextual Bandits with Graph-Laplacian Regularization—0
Show:102550
← PrevPage 22 of 27Next →

No leaderboard results yet.