SOTAVerified

Multi-Armed Bandits

Multi-armed bandits refer to a task where a fixed amount of resources must be allocated between competing resources that maximizes expected gain. Typically these problems involve an exploration/exploitation trade-off.

( Image credit: Microsoft Research )

Papers

Showing 601–650 of 1262 papers

TitleStatusHype
Multi-armed Bandits for Link Configuration in Millimeter-wave Networks—0
Adaptive Experimentation with Delayed Binary FeedbackCode0
Scalable Decision-Focused Learning in Restless Multi-Armed Bandits with Application to Maternal and Child Health—0
Efficient Algorithms for Learning to Control Bandits with Unobserved Contexts—0
Context Uncertainty in Contextual Bandits with Applications to Recommender Systems—0
Optimal Regret Is Achievable with Bounded Approximate Inference Error: An Enhanced Bayesian Upper Confidence Bound FrameworkCode0
Evaluating Deep Vs. Wide & Deep Learners As Contextual Bandits For Personalized Email Promo RecommendationsCode0
Neural Collaborative Filtering Bandits via Meta Learning—0
Coordinated Attacks against Contextual Bandits: Fundamental Limits and Defense Mechanisms—0
Networked Restless Multi-Armed Bandits for Mobile Interventions—0
Top-K Ranking Deep Contextual Bandits for Information Selection Systems—0
Adaptive Best-of-Both-Worlds Algorithm for Heavy-Tailed Multi-Armed Bandits—0
Learning Neural Contextual Bandits Through Perturbed Rewards—0
Occupancy Information Ratio: Infinite-Horizon, Information-Directed, Parameterized Policy Search—0
Semantic Parsing for Planning Goals as Constrained Combinatorial Contextual Bandits—0
Contextual Bandits for Advertising Campaigns: A Diffusion-Model Independent Approach (Extended Version)—0
Modelling Cournot Games as Multi-agent Multi-armed Bandits—0
Off-Policy Evaluation Using Information Borrowing and Context-Based SwitchingCode0
Stochastic differential equations for limiting description of UCB rule for Gaussian multi-armed bandits—0
Safe Linear Leveling Bandits—0
Privacy Amplification via Shuffling for Linear Contextual Bandits—0
Efficient Action Poisoning Attacks on Linear Contextual Bandits—0
Best Arm Identification under Additive Transfer Bandits—0
Contextual Bandit Applications in Customer Support Bot—0
On Submodular Contextual Bandits—0
Optimal Algorithms for Stochastic Contextual Preference Bandits—0
Identification of the Generalized Condorcet Winner in Multi-dueling BanditsCode0
Asymptotically Best Causal Effect Identification with Multi-Armed Bandits—0
Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningCode0
Bandits with Knapsacks beyond the Worst Case—0
Multi-Armed Bandits with Bounded Arm-Memory: Near-Optimal Guarantees for Best-Arm Identification and Regret Minimization—0
Online Fair Revenue Maximizing Cake Division with Non-Contiguous Pieces in Adversarial Bandits—0
Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationCode1
Decentralized Upper Confidence Bound Algorithms for Homogeneous Multi-Agent Multi-Armed Bandits—0
Offline Contextual Bandits for Wireless Network Optimization—0
An Instance-Dependent Analysis for the Cooperative Multi-Player Multi-Armed Bandit—0
Universal and data-adaptive algorithms for model selection in linear contextual bandits—0
Empirical analysis of representation learning and exploration in neural kernel banditsCode0
Privacy-Preserving Communication-Efficient Federated Multi-Armed Bandits—0
Bandits Don’t Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits—0
Decentralized Cooperative Reinforcement Learning with Hierarchical Information Structure—0
(Almost) Free Incentivized Exploration from Decentralized Learning AgentsCode0
Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and GeneralizationCode0
Federated Linear Contextual Bandits—0
The Pareto Frontier of model selection for general Contextual Bandits—0
Linear Contextual Bandits with Adversarial Corruptions—0
Analysis of Thompson Sampling for Partially Observable Contextual Multi-Armed Bandits—0
Towards the D-Optimal Online Experiment Design for Recommender SelectionCode0
Dynamic pricing and assortment under a contextual MNL demand—0
Stateful Offline Contextual Policy Evaluation and Learning—0
Show:102550
← PrevPage 13 of 26Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1NeuralLinear FullPosterior-MRCumulative regret1.92—Unverified
2Linear FullPosterior-MRCumulative regret1.82—Unverified