Papers
Topics
Authors
Recent
Search
2000 character limit reached

Bounded Memory Adversarial Bandits with Composite Anonymous Delayed Feedback

Published 27 Apr 2022 in cs.LG and stat.ML | (2204.12764v2)

Abstract: We study the adversarial bandit problem with composite anonymous delayed feedback. In this setting, losses of an action are split into dd components, spreading over consecutive rounds after the action is chosen. And in each round, the algorithm observes the aggregation of losses that come from the latest dd rounds. Previous works focus on oblivious adversarial setting, while we investigate the harder non-oblivious setting. We show non-oblivious setting incurs Ω(T)\Omega(T) pseudo regret even when the loss sequence is bounded memory. However, we propose a wrapper algorithm which enjoys o(T)o(T) policy regret on many adversarial bandit problems with the assumption that the loss sequence is bounded memory. Especially, for KK-armed bandit and bandit convex optimization, we have O(T<sup>2/3)\mathcal{O}(T<sup>{2/3}) policy regret bound. We also prove a matching lower bound for KK-armed bandit. Our lower bound works even when the loss sequence is oblivious but the delay is non-oblivious. It answers the open problem proposed in \cite{wang2021adaptive}, showing that non-oblivious delay is enough to incur Ω~(T<sup>2/3)\tilde{\Omega}(T<sup>{2/3}) regret.

Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.