SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 76–100 of 655 papers

TitleStatusHype
Fast Change Identification in Multi-Play Bandits and its Applications in Wireless Networks—0
A Bayesian Choice Model for Eliminating Feedback Loops—0
A Practical Method for Solving Contextual Bandit Problems Using Decision Trees—0
A Provably Efficient Model-Free Posterior Sampling Method for Episodic Reinforcement Learning—0
Efficiently Tackling Million-Dimensional Multiobjective Problems: A Direction Sampling and Fine-Tuning Approach—0
A Reinforcement Learning based Reset Policy for CDCL SAT Solvers—0
A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems—0
A Reliability-aware Multi-armed Bandit Approach to Learn and Select Users in Demand Response—0
A resource-constrained stochastic scheduling algorithm for homeless street outreach and gleaning edible food—0
A sequential Monte Carlo approach to Thompson sampling for Bayesian optimization—0
A Simple and Optimal Policy Design with Safety against Heavy-Tailed Risk for Stochastic Bandits—0
A study of Thompson Sampling with Parameter h—0
Asymptotically Optimal Algorithms for Budgeted Multiple Play Bandits—0
Asymptotically Optimal Bandits under Weighted Information—0
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget—0
The Choice of Noninformative Priors for Thompson Sampling in Multiparameter Bandit Models—0
Asymptotic Convergence of Thompson Sampling—0
Asymptotic Performance of Thompson Sampling in the Batched Multi-Armed Bandits—0
Aging Bandits: Regret Analysis and Order-Optimal Learning Algorithm for Wireless Networks with Stochastic Arrivals—0
Apple Tasting Revisited: Bayesian Approaches to Partially Monitored Online Binary Classification—0
Asynchronous Multi Agent Active Search—0
Algorithms for Adaptive Experiments that Trade-off Statistical Analysis with Reward: Combining Uniform Random Assignment and Reward Maximization—0
An Unbiased Data Collection and Content Exploitation/Exploration Strategy for Personalization—0
Augmented RBMLE-UCB Approach for Adaptive Control of Linear Quadratic Systems—0
Adaptive Sensor Placement for Continuous Spaces—0
Show:102550
← PrevPage 4 of 27Next →

No leaderboard results yet.