SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 1–10 of 655 papers

TitleStatusHype
Robust Policy Switching for Antifragile Reinforcement Learning for UAV Deconfliction in Adversarial Environments—0
Context Attribution with Multi-Armed Bandit Optimization—0
Adaptive Data Augmentation for Thompson Sampling—0
Bayesian Optimization with Inexact Acquisition: Is Random Grid Search Sufficient?—0
Efficient kernelized bandit algorithms via exploration distributions—0
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget—0
Simplifying Bayesian Optimization Via In-Context Direct Optimum Sampling—0
Thompson Sampling in Online RLHF with General Function Approximation—0
Stable Thompson Sampling: Valid Inference via Variance Inflation—0
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection—0
Show:102550
← PrevPage 1 of 66Next →

No leaderboard results yet.