Papers
Topics
Authors
Recent
Search
2000 character limit reached

Risk-averse Contextual Multi-armed Bandit Problem with Linear Payoffs

Published 24 Jun 2022 in cs.LG, cs.IT, math.IT, and stat.ML | (2206.12463v1)

Abstract: In this paper we consider the contextual multi-armed bandit problem for linear payoffs under a risk-averse criterion. At each round, contexts are revealed for each arm, and the decision maker chooses one arm to pull and receives the corresponding reward. In particular, we consider mean-variance as the risk criterion, and the best arm is the one with the largest mean-variance reward. We apply the Thompson Sampling algorithm for the disjoint model, and provide a comprehensive regret analysis for a variant of the proposed algorithm. For TT rounds, KK actions, and dd-dimensional feature vectors, we prove a regret bound of O((1+ρ+1ρ)dlnTlnKδdKT<sup>1+2ϵ</sup>lnKδ1ϵ)O((1+\rho+\frac{1}{\rho}) d\ln T \ln \frac{K}{\delta}\sqrt{d K T<sup>{1+2\epsilon}</sup> \ln \frac{K}{\delta} \frac{1}{\epsilon}}) that holds with probability 1δ1-\delta under the mean-variance criterion with risk tolerance ρ\rho, for any $0&lt;\epsilon&lt;\frac{1}{2}$, $0&lt;\delta&lt;1$. The empirical performance of our proposed algorithms is demonstrated via a portfolio selection problem.

Authors (3)
Citations (4)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.