SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 601–650 of 655 papers

TitleStatusHype
Stacked Thompson BanditsCode0
Thompson Sampling For Stochastic Bandits with Graph Feedback—0
Estimating Quality in Multi-Objective Bandits Optimization—0
Exploration for Multi-task Reinforcement Learning with Deep Generative Models—0
Nonparametric General Reinforcement Learning—0
Linear Thompson Sampling Revisited—0
Unimodal Thompson Sampling for Graph-Structured Arms—0
The End of Optimism? An Asymptotic Analysis of Finite-Armed Linear Bandits—0
A Formal Solution to the Grain of Truth Problem—0
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems—0
Human collective intelligence as distributed Bayesian inference—0
Asymptotically Optimal Algorithms for Budgeted Multiple Play Bandits—0
Online Algorithms For Parameter Mean And Variance Estimation In Dynamic Regression Models—0
Linear Bandit algorithms using the Bootstrap—0
Double Thompson Sampling for Dueling BanditsCode0
An Unbiased Data Collection and Content Exploitation/Exploration Strategy for Personalization—0
A sequential Monte Carlo approach to Thompson sampling for Bayesian optimization—0
Optimal Recommendation to Users that React: Online Learning for a Class of POMDPs—0
Cascading Bandits for Large-Scale Recommendation ProblemsCode0
Simple Bayesian Algorithms for Best Arm Identification—0
Thompson Sampling is Asymptotically Optimal in General Environments—0
Convolutional Monte Carlo Rollouts in Go—0
Efficient Thompson Sampling for Online Matrix-Factorization Recommendation—0
Regret Analysis of the Finite-Horizon Gittins Index Strategy for Multi-Armed Bandits—0
TSEB: More Efficient Thompson Sampling for Policy Learning—0
Incentivizing Exploration In Reinforcement Learning With Deep Predictive ModelsCode0
Bootstrapped Thompson Sampling and Deep Exploration—0
On the Prior Sensitivity of Thompson Sampling—0
Optimal Regret Analysis of Thompson Sampling in Stochastic Multi-armed Bandit Problem with Multiple PlaysCode0
Belief Flows of Robust Online Learning—0
Thompson Sampling for Budgeted Multi-armed Bandits—0
Evaluation of Explore-Exploit Policies in Multi-result Ranking Systems—0
A Note on Information-Directed Sampling and Thompson Sampling—0
Bandit Convex Optimization: sqrtT Regret in One Dimension—0
Thompson sampling with the online bootstrap—0
Freshness-Aware Thompson Sampling—0
Towards Optimal Algorithms for Prediction with Expert Advice—0
Thompson Sampling for Learning Parameterized Markov Decision Processes—0
Efficient Learning in Large-Scale Combinatorial Semi-Bandits—0
An Information-Theoretic Analysis of Thompson Sampling—0
Better Optimism By Bayes: Adaptive Planning with Rich Models—0
Bayesian Mixture Modelling and Inference based Thompson Sampling in Monte-Carlo Tree Search—0
Eluder Dimension and the Sample Complexity of Optimistic Exploration—0
Thompson Sampling for Complex Bandit Problems—0
Thompson Sampling for Online Learning with Linear Experts—0
Generalized Thompson Sampling for Contextual Bandits—0
Thompson Sampling in Dynamic Systems for Contextual Bandit Problems—0
Thompson Sampling for 1-Dimensional Exponential Family Bandits—0
Cover Tree Bayesian Reinforcement Learning—0
Prior-free and prior-dependent regret bounds for Thompson Sampling—0
Show:102550
← PrevPage 13 of 14Next →

No leaderboard results yet.