Papers
Topics
Authors
Recent
Search
2000 character limit reached

Adversarial Multi-dueling Bandits

Published 18 Jun 2024 in cs.LG | (2406.12475v2)

Abstract: We introduce the problem of regret minimization in adversarial multi-dueling bandits. While adversarial preferences have been studied in dueling bandits, they have not been explored in multi-dueling bandits. In this setting, the learner is required to select m2m \geq 2 arms at each round and observes as feedback the identity of the most preferred arm which is based on an arbitrary preference matrix chosen obliviously. We introduce a novel algorithm, MiDEX (Multi Dueling EXP3), to learn from such preference feedback that is assumed to be generated from a pairwise-subset choice model. We prove that the expected cumulative TT-round regret of MiDEX compared to a Borda-winner from a set of KK arms is upper bounded by O((KlogK)<sup>1/3</sup>T<sup>2/3)O((K \log K)<sup>{1/3}</sup> T<sup>{2/3}). Moreover, we prove a lower bound of Ω(K<sup>1/3</sup>T<sup>2/3)\Omega(K<sup>{1/3}</sup> T<sup>{2/3}) for the expected regret in this setting which demonstrates that our proposed algorithm is near-optimal.

Authors (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.