SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 501–550 of 655 papers

TitleStatusHype
Thompson Sampling with Virtual Helping Agents—0
Time-Sensitive Bandit Learning and Satisficing Thompson Sampling—0
Top Two Algorithms Revisited—0
Towards Optimal Algorithms for Prediction with Expert Advice—0
Towards Scalable and Robust Structured Bandits: A Meta-Learning Framework—0
Tree Ensembles for Contextual Bandits—0
Truthful mechanisms for linear bandit games with private contexts—0
TSEB: More Efficient Thompson Sampling for Policy Learning—0
TSEC: a framework for online experimentation under experimental constraints—0
TS-UCB: Improving on Thompson Sampling With Little to No Additional Computation—0
Two-Stage Resource Allocation in Reconfigurable Intelligent Surface Assisted Hybrid Networks via Multi-Player Bandits—0
Uncertainty-Aware Search and Value Models: Mitigating Search Scaling Flaws in LLMs—0
Understanding the Training and Generalization of Pretrained Transformer for Sequential Decision Making—0
Reinforcement Learning in Credit Scoring and Underwriting—0
Unimodal Thompson Sampling for Graph-Structured Arms—0
Using Adaptive Experiments to Rapidly Help Students—0
Variable Selection via Thompson Sampling—0
Variational Bayesian Optimistic Sampling—0
WAPTS: A Weighted Allocation Probability Adjusted Thompson Sampling Algorithm for High-Dimensional and Sparse Experiment Settings—0
When and Whom to Collaborate with in a Changing Environment: A Collaborative Dynamic Bandit Solution—0
When and why randomised exploration works (in linear bandits)—0
When Combinatorial Thompson Sampling meets Approximation Regret—0
Practical Batch Bayesian Sampling Algorithms for Online Adaptive Traffic Experimentation—0
Zero-Inflated Bandits—0
A Bandit Approach to Online Pricing for Heterogeneous Edge Resource Allocation—0
A Batched Multi-Armed Bandit Approach to News Headline Testing—0
Context in Public Health for Underserved Communities: A Bayesian Approach to Online Restless Bandits—0
A Bayesian Choice Model for Eliminating Feedback Loops—0
Accelerating Grasp Exploration by Leveraging Learned Priors—0
A Change-Detection Based Thompson Sampling Framework for Non-Stationary Bandits—0
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling—0
A Closer Look at the Worst-case Behavior of Multi-armed Bandit Algorithms—0
A Combinatorial Semi-Bandit Approach to Charging Station Selection for Electric Vehicles—0
A Contextual Combinatorial Semi-Bandit Approach to Network Bottleneck Identification—0
A Copula approach for hyperparameter transfer learning—0
A Quantile-based Approach for Hyperparameter Transfer Learning—0
Fast Change Identification in Multi-Play Bandits and its Applications in Wireless Networks—0
Active Reinforcement Learning with Monte-Carlo Tree Search—0
Active Search for High Recall: a Non-Stationary Extension of Thompson Sampling—0
AdaptEx: A Self-Service Contextual Bandit Platform—0
Adaptive Combinatorial Allocation—0
Adaptive Data Augmentation for Thompson Sampling—0
Adaptive Experimentation at Scale: A Computational Framework for Flexible Batches—0
Adaptive Exploration-Exploitation Tradeoff for Opportunistic Bandits—0
Adaptive Gating for Single-Photon 3D Imaging—0
Adaptive Grey-Box Fuzz-Testing with Thompson Sampling—0
Adaptively Learning to Select-Rank in Online Platforms—0
Adaptively Optimize Content Recommendation Using Multi Armed Bandit Algorithms in E-commerce—0
Adaptive Model Selection Framework: An Application to Airline Pricing—0
Adaptive Operator Selection Based on Dynamic Thompson Sampling for MOEA/D—0
Show:102550
← PrevPage 11 of 14Next →

No leaderboard results yet.