Evaluating policies for generalized bandits via a notion of duality-Reference-Cited by-同舟云学术

Evaluating policies for generalized bandits via a notion of duality

Published:2000-06 Issue:02 Volume:37 Page:540-546
ISSN:0021-9002
Container-title:Journal of Applied Probability
language:en
Short-container-title:J. Appl. Probab.

Author:

Crosbie J. H.,Glazebrook K. D.

Abstract

Nash's generalization of Gittins’ classic index result to so-called generalized bandit problems (GBPs) in which returns are dependent on the states of all arms (not only the one which is pulled) has proved important for applications. The index theory for special cases of this model in which all indices are positive is straightforward. However, this is not a natural restriction in practice. An earlier proposal for the general case did not yield satisfactory index-based suboptimality bounds for policies — a central feature of classical Gittins index theory. We develop such bounds via a notion of duality for GBPs which is of independent interest. The index which emerges naturally from this analysis is the reciprocal of the one proposed by Nash.

Publisher

Cambridge University Press (CUP)

Subject

Statistics, Probability and Uncertainty,General Mathematics,Statistics and Probability

Reference14 articles.

1. On scheduling influential stochastic tasks on a single machine

2. Evaluating strategies for generalized bandit problems

3. On the evaluation of suboptimal strategies for families of alternative bandit processes

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Optimal index shooting policy for layered missile defense system;Journal of Systems Engineering and Electronics;2020-01-31

2. Index Policies for Shooting Problems;Operations Research;2007-08