Department of Mathematics
 Search | Help | Login | pdf version | printable version

Math @ Duke





.......................

.......................


Publications [#336644] of Vahid Tarokh

Papers Published

  1. Shahrampour, S; Tarokh, V, Nonlinear sequential accepts and rejects for identification of top arms in stochastic bandits, 55th Annual Allerton Conference on Communication, Control, and Computing, Allerton 2017, vol. 2018-January (July, 2017), pp. 228-235, IEEE [doi]
    (last updated on 2023/06/01)

    Abstract:
    We address the M-best-Arm identification problem in multi-Armed bandits. A player has a limited budget to explore K arms (M < K), and once pulled, each arm yields a reward drawn (independently) from a fixed, unknown distribution. The goal is to find the top M arms in the sense of expected reward. We develop an algorithm which proceeds in rounds to deactivate arms iteratively. At each round, the budget is divided by a nonlinear function of remaining arms, and the arms are pulled correspondingly. Based on a decision rule, the deactivated arm at each round may be accepted or rejected. The algorithm outputs the accepted arms that should ideally be the top M arms. We characterize the decay rate of the misidentification probability and establish that the nonlinear budget allocation proves to be useful for different problem environments (described by the number of competitive arms). We provide comprehensive numerical experiments showing that our algorithm outperforms the state-of-The-Art using suitable nonlinearity.

 

dept@math.duke.edu
ph: 919.660.2800
fax: 919.660.2821

Mathematics Department
Duke University, Box 90320
Durham, NC 27708-0320