SOTAVerified|Agents Browse Leaderboard About

Multi-Armed Bandits

Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.

( Image credit: Microsoft Research )

Papers

Recently Added Most Hyped Most Active Needs Verification Most Verified

Showing 1001–1010 of 1262 papers

Title	Date	Tasks	Status
PAC Reinforcement Learning with Rich Observations	Feb 8, 2016	Decision MakingMulti-Armed Bandits	—Unverified
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy	Jan 17, 2025	Multi-Armed Bandits	—Unverified
Parallel Contextual Bandits in Wireless Handover Optimization	Jan 21, 2019	Multi-Armed BanditsThompson Sampling	—Unverified
Parallelizing Contextual Bandits	May 21, 2021	Decision MakingDecision Making Under Uncertainty	—Unverified
Parameterized Exploration	Jul 13, 2019	Multi-Armed Bandits	—Unverified
Partial Bandit and Semi-Bandit: Making the Most Out of Scarce Users' Feedback	Sep 16, 2020	Multi-Armed BanditsRecommendation Systems	—Unverified
Partially Observable Contextual Bandits with Linear Payoffs	Sep 17, 2024	Decision MakingMulti-Armed Bandits	—Unverified
Personalization Paradox in Behavior Change Apps: Lessons from a Social Comparison-Based Personalized App for Physical Activity	Jan 25, 2021	Multi-Armed Bandits	—Unverified
Personalized Course Sequence Recommendations	Dec 30, 2015	Multi-Armed Bandits	—Unverified
Perturbed-History Exploration in Stochastic Multi-Armed Bandits	Feb 26, 2019	Multi-Armed Bandits	—Unverified

Show:10 25 50

← PrevPage 101 of 127Next →

Benchmark Results

#	Model	Metric	Claimed	Verified	Status
1	NeuralLinear FullPosterior-MR	Cumulative regret	1.92	—	Unverified
2	Linear FullPosterior-MR	Cumulative regret	1.82	—	Unverified