SOTAVerified

Multi-Armed Bandits

Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.

( Image credit: Microsoft Research )

Papers

Showing 151–175 of 1262 papers

TitleStatusHype
Bandit Regret Scaling with the Effective Loss Range—0
Bandits Don't Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits—0
Bandits Don’t Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits—0
Bandits for Learning to Explain from Explanations—0
Bandits meet Computer Architecture: Designing a Smartly-allocated Cache—0
Bandit Social Learning: Exploration under Myopic Behavior—0
Bandits Warm-up Cold Recommender Systems—0
Preferences Evolve And So Should Your Bandits: Bandits with Evolving States for Online Platforms—0
Bandits with Knapsacks beyond the Worst Case—0
Bandits with Partially Observable Confounded Data—0
Bandits with Temporal Stochastic Constraints—0
Banker Online Mirror Descent—0
Banker Online Mirror Descent: A Universal Approach for Delayed Online Bandit Learning—0
Batched Bandits with Crowd Externalities—0
Batched Coarse Ranking in Multi-Armed Bandits—0
Almost Optimal Batch-Regret Tradeoff for Batch Linear Contextual Bandits—0
Regret Bounds for Batched Bandits—0
Batched Nonparametric Bandits via k-Nearest Neighbor UCB—0
A Gang of Bandits—0
Batched Online Contextual Sparse Bandits with Sequential Inclusion of Features—0
Batched Thompson Sampling—0
Batched Thompson Sampling for Multi-Armed Bandits—0
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits—0
Towards Bayesian Data Selection—0
Balanced off-policy evaluation in general action spaces—0
Show:102550
← PrevPage 7 of 51Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NeuralLinear FullPosterior-MRCumulative regret1.92—Unverified
2Linear FullPosterior-MRCumulative regret1.82—Unverified