Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stochastic Bandits with Linear Constraints

Published 17 Jun 2020 in cs.LG and stat.ML | (2006.10185v1)

Abstract: We study a constrained contextual linear bandit setting, where the goal of the agent is to produce a sequence of policies, whose expected cumulative reward over the course of TT rounds is maximum, and each has an expected cost below a certain threshold τ\tau. We propose an upper-confidence bound algorithm for this problem, called optimistic pessimistic linear bandit (OPLB), and prove an O~(dTτ−c0)\widetilde{\mathcal{O}}(\frac{d\sqrt{T}}{\tau-c_0}) bound on its TT-round regret, where the denominator is the difference between the constraint threshold and the cost of a known feasible action. We further specialize our results to multi-armed bandits and propose a computationally efficient algorithm for this setting. We prove a regret bound of O~(KTτ−c0)\widetilde{\mathcal{O}}(\frac{\sqrt{KT}}{\tau - c_0}) for this algorithm in KK-armed bandits, which is a K\sqrt{K} improvement over the regret bound we obtain by simply casting multi-armed bandits as an instance of contextual linear bandits and using the regret bound of OPLB. We also prove a lower-bound for the problem studied in the paper and provide simulations to validate our theoretical results.

Citations (63)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.