SOTAVerified|Agents Browse Leaderboard About

Multi-Armed Bandits

Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.

( Image credit: Microsoft Research )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 491–500 of 1262 papers

Title	Date	Tasks	Status
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits	Sep 13, 2024	Multi-Armed BanditsReinforcement Learning (RL)	—Unverified
FedMABA: Towards Fair Federated Learning through Multi-Armed Bandits Allocation	Oct 26, 2024	FairnessFederated Learning	—Unverified
Feel-Good Thompson Sampling for Contextual Bandits and Reinforcement Learning	Oct 2, 2021	Multi-Armed Banditsregression	—Unverified
Feel-Good Thompson Sampling for Contextual Dueling Bandits	Apr 9, 2024	Decision MakingMulti-Armed Bandits	—Unverified
Analysis of Thompson Sampling for Partially Observable Contextual Multi-Armed Bandits	Oct 23, 2021	Decision MakingMulti-Armed Bandits	—Unverified
Fighting Contextual Bandits with Stochastic Smoothing	Oct 11, 2018	Multi-Armed Bandits	—Unverified
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy	Jan 24, 2025	Decision MakingMulti-Armed Bandits	—Unverified
Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child Health	Feb 2, 2022	Multi-Armed BanditsScheduling	—Unverified
Finding the bandit in a graph: Sequential search-and-stop	Jun 6, 2018	Multi-Armed Bandits	—Unverified
Batched Thompson Sampling for Multi-Armed Bandits	Aug 15, 2021	Multi-Armed BanditsThompson Sampling	—Unverified

Show:10 25 50

← PrevPage 50 of 127Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	NeuralLinear FullPosterior-MR	Cumulative regret	1.92	—	Unverified
2	Linear FullPosterior-MR	Cumulative regret	1.82	—	Unverified