SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 326350 of 655 papers

TitleStatusHype
On Dynamic Pricing with Covariates0
On Efficiency in Hierarchical Reinforcement Learning0
On Improved Regret Bounds In Bayesian Optimization with Gaussian Noise0
On Kernelized Multi-Armed Bandits with Constraints0
On learning Whittle index policy for restless bandits with scalable regret0
Online Algorithms For Parameter Mean And Variance Estimation In Dynamic Regression Models0
Online Continuous Hyperparameter Optimization for Generalized Linear Contextual Bandits0
Online Causal Inference for Advertising in Real-Time Bidding Auctions0
Online Learning and Distributed Control for Residential Demand Response0
Online Learning-based Waveform Selection for Improved Vehicle Recognition in Automotive Radar0
Online Learning of Energy Consumption for Navigation of Electric Vehicles0
Online Learning of Network Bottlenecks via Minimax Paths0
Online Residential Demand Response via Contextual Multi-Armed Bandits0
Only Pay for What Is Uncertain: Variance-Adaptive Thompson Sampling0
On Multi-Armed Bandit Designs for Dose-Finding Clinical Trials0
On Online Learning in Kernelized Markov Decision Processes0
On The Differential Privacy of Thompson Sampling With Gaussian Prior0
On the Importance of Uncertainty in Decision-Making with Large Language Models0
On the Performance of Thompson Sampling on Logistic Bandits0
On the Prior Sensitivity of Thompson Sampling0
On Thompson Sampling for Smoother-than-Lipschitz Bandits0
On Thompson Sampling with Langevin Algorithms0
On Frequentist Regret of Linear Thompson Sampling0
Near-Optimal Algorithms for Differentially Private Online Learning in a Stochastic Environment0
Optimal Exploration is no harder than Thompson Sampling0
Show:102550
← PrevPage 14 of 27Next →

No leaderboard results yet.