Papers
Topics
Authors
Recent
Search
2000 character limit reached

Thompson Sampling for Combinatorial Semi-bandits with Sleeping Arms and Long-Term Fairness Constraints

Published 14 May 2020 in cs.LG and stat.ML | (2005.06725v1)

Abstract: We study the combinatorial sleeping multi-armed semi-bandit problem with long-term fairness constraints~(CSMAB-F). To address the problem, we adopt Thompson Sampling~(TS) to maximize the total rewards and use virtual queue techniques to handle the fairness constraints, and design an algorithm called \emph{TS with beta priors and Bernoulli likelihoods for CSMAB-F~(TSCSF-B)}. Further, we prove TSCSF-B can satisfy the fairness constraints, and the time-averaged regret is upper bounded by N2η+O(mNTlnTT)\frac{N}{2\eta} + O\left(\frac{\sqrt{mNT\ln T}}{T}\right), where NN is the total number of arms, mm is the maximum number of arms that can be pulled simultaneously in each round~(the cardinality constraint) and η\eta is the parameter trading off fairness for rewards. By relaxing the fairness constraints (i.e., let η\eta \rightarrow \infty), the bound boils down to the first problem-independent bound of TS algorithms for combinatorial sleeping multi-armed semi-bandit problems. Finally, we perform numerical experiments and use a high-rating movie recommendation application to show the effectiveness and efficiency of the proposed algorithm.

Citations (9)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.