SOTAVerified

Thompson Sampling

Thompson sampling, named after William R. Thompson, is a heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.

Papers

Showing 101–150 of 655 papers

TitleStatusHype
A Multi-Armed Bandit to Smartly Select a Training Set from Big Medical Data—0
Adaptive Combinatorial Allocation—0
Automatic Ensemble Learning for Online Influence Maximization—0
AutoSeM: Automatic Task Selection and Mixing in Multi-Task Learning—0
Bag of Policies for Distributional Deep Exploration—0
BanditCAT and AutoIRT: Machine Learning Approaches to Computerized Adaptive Testing and Item Calibration—0
Bandit Change-Point Detection for Real-Time Monitoring High-Dimensional Data Under Sampling Control—0
Bandit Convex Optimization: sqrtT Regret in One Dimension—0
Bandit Learning for Diversified Interactive Recommendation—0
Adaptive Rate of Convergence of Thompson Sampling for Gaussian Process Optimization—0
Bandit Models of Human Behavior: Reward Processing in Mental Disorders—0
Bandit Policies for Reliable Cellular Network Handovers in Extreme Mobility—0
Bandits Under The Influence (Extended Version)—0
Bandit Theory and Thompson Sampling-Guided Directed Evolution for Sequence Optimization—0
Batch Bayesian Optimization for Replicable Experimental Design—0
Adaptive Sensor Placement for Continuous Spaces—0
Batched Thompson Sampling—0
Batched Thompson Sampling for Multi-Armed Bandits—0
An Arm-Wise Randomization Approach to Combinatorial Linear Semi-Bandits—0
Bayesian Bandit Algorithms with Approximate Inference in Stochastic Linear Bandits—0
An Efficient Algorithm For Generalized Linear Bandit: Online Stochastic Gradient Descent and Thompson Sampling—0
Bayesian Best-Arm Identification for Selecting Influenza Mitigation Strategies—0
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff—0
Bayesian decision-making under misspecified priors with applications to meta-learning—0
Bayesian-Guided Generation of Synthetic Microbiomes with Minimized Pathogenicity—0
Bayesian Learning of Optimal Policies in Markov Decision Processes with Countably Infinite State-Space—0
Adaptive Operator Selection Based on Dynamic Thompson Sampling for MOEA/D—0
Tsallis-INF: An Optimal Algorithm for Stochastic and Adversarial Bandits—0
A Quantile-based Approach for Hyperparameter Transfer Learning—0
Bayesian Analysis of Combinatorial Gaussian Process Bandits—0
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing—0
A Nonparametric Contextual Bandit with Arm-level Eligibility Control for Customer Service Routing—0
An Online Learning Framework for Energy-Efficient Navigation of Electric Vehicles—0
Adaptive Model Selection Framework: An Application to Airline Pricing—0
Belief Flows of Robust Online Learning—0
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems—0
An Information-Theoretic Analysis of Thompson Sampling with Infinite Action Spaces—0
BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems—0
Best Arm Identification in Batched Multi-armed Bandit Problems—0
Active RLHF via Best Policy Learning from Trajectory Preference Feedback—0
Better Optimism By Bayes: Adaptive Planning with Rich Models—0
Blind Exploration and Exploitation of Stochastic Experts—0
Bootstrapped Thompson Sampling and Deep Exploration—0
BOTS: Batch Bayesian Optimization of Extended Thompson Sampling for Severely Episode-Limited RL Settings—0
Calibrated Fairness in Bandits—0
A Note on Information-Directed Sampling and Thompson Sampling—0
An Unbiased Data Collection and Content Exploitation/Exploration Strategy for Personalization—0
Causal Bandits without prior knowledge using separating sets—0
Chained Information-Theoretic bounds and Tight Regret Rate for Linear Bandit Problems—0
Bayesian Quantile and Expectile Optimisation—0
Show:102550
← PrevPage 3 of 14Next →

No leaderboard results yet.