Papers
Topics
Authors
Recent
Search
2000 character limit reached

Pure Exploration with Structured Preference Feedback

Published 12 Apr 2021 in cs.LG | (2104.05294v1)

Abstract: We consider the problem of pure exploration with subset-wise preference feedback, which contains NN arms with features. The learner is allowed to query subsets of size KK and receives feedback in the form of a noisy winner. The goal of the learner is to identify the best arm efficiently using as few queries as possible. This setting is relevant in various online decision-making scenarios involving human feedback such as online retailing, streaming services, news feed, and online advertising; since it is easier and more reliable for people to choose a preferred item from a subset than to assign a likability score to an item in isolation. To the best of our knowledge, this is the first work that considers the subset-wise preference feedback model in a structured setting, which allows for potentially infinite set of arms. We present two algorithms that guarantee the detection of the best-arm in O~(d<sup>2K</sup>Δ<sup>2)\tilde{O} (\frac{d<sup>2}{K</sup> \Delta<sup>2}) samples with probability at least 1δ1 - \delta, where dd is the dimension of the arm-features and Δ\Delta is the appropriate notion of utility gap among the arms. We also derive an instance-dependent lower bound of Ω(dΔ<sup>2</sup>log1δ)\Omega(\frac{d}{\Delta<sup>2}</sup> \log \frac{1}{\delta}) which matches our upper bound on a worst-case instance. Finally, we run extensive experiments to corroborate our theoretical findings, and observe that our adaptive algorithm stops and requires up to 12x fewer samples than a non-adaptive algorithm.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.