Gaussian One-Armed Bandit Problem

We consider the multi-armed bandit problem with penalties for switching that include setup delays and costs, extending the former results of the author for the special case with no switching delays. A priority index for projects with setup delays that characterizes, in part, optimal policies was introduced by Asawa and Teneketzis in 1996, yet without giving a means of computing it. We present a fast two-stage index computing method, which computes the continuation index (which applies when the project has been set up) in a first stage and certain extra quantities with cubic (arithmetic-operation) complexity in the number of project states and then computes the switching index (which applies when the project is not set up), in a second stage, with quadratic complexity. The approach is based on new methodological advances on restless bandit indexation, which are introduced and deployed herein, being motivated by the limitations of previous results, exploiting the fact that the aforementioned index is the Whittle index of the project in its restless reformulation. A numerical study demonstrates substantial runtime speed-ups of the new two-stage index algorithm versus a general one-stage Whittle index algorithm. The study further gives evidence that, in a multi-project setting, the index policy is consistently nearly optimal.

Download Full-text

On transforming an index for generalised bandit problems

Journal of Applied Probability ◽

10.2307/3214927 ◽

1995 ◽

Vol 32 (1) ◽

pp. 168-182 ◽

Cited By ~ 4

Author(s):

K. D. Glazebrook ◽

S. Greatrix

Keyword(s):

Dynamic Programming ◽

Policy Evaluation ◽

Gittins Index ◽

Bandit Problem ◽

Bandit Problems ◽

Index Policies

Nash (1980) demonstrated that index policies are optimal for a class of generalised bandit problem. A transform of the index concerned has many of the attributes of the Gittins index. The transformed index is positive-valued, with maximal values yielding optimal actions. It may be characterised as the value of a restart problem and is hence computable via dynamic programming methodologies. The transformed index can also be used in procedures for policy evaluation.

Download Full-text

Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-Part I: I.I.D. rewards

IEEE Transactions on Automatic Control ◽

10.1109/tac.1987.1104491 ◽

1987 ◽

Vol 32 (11) ◽

pp. 968-976 ◽

Cited By ~ 102

Author(s):

V. Anantharam ◽

P. Varaiya ◽

J. Walrand

Keyword(s):

Efficient Allocation ◽

Bandit Problem ◽

Allocation Rules ◽

Asymptotically Efficient ◽

Multiarmed Bandit

Download Full-text

Bayesian Learning in an Unstable World: Experimental Evidence Based on the Bandit Problem

SSRN Electronic Journal ◽

10.2139/ssrn.1628657 ◽

2012 ◽

Cited By ~ 2

Author(s):

Elise Payzan-LeNestour

Keyword(s):

Experimental Evidence ◽

Bayesian Learning ◽

Evidence Based ◽

Bandit Problem

Download Full-text

On a Two-armed Bandit Problem with both Continuous and Impulse Actions and Discounted Rewards

Seminar on Stochastic Processes, 1992 ◽

10.1007/978-1-4612-0339-1_13 ◽

1993 ◽

pp. 249-266

Author(s):

A. A. Yushkevich

Keyword(s):

Bandit Problem ◽

Discounted Rewards

Download Full-text

Exact solution of the Bellman equation for a β-discounted reward in a two-armed bandit with switching arms

Journal of Applied Mathematics and Stochastic Analysis ◽

10.1155/s1048953399000155 ◽

1999 ◽

Vol 12 (2) ◽

pp. 151-160 ◽

Cited By ~ 1

Author(s):

Doncho S. Donchev

Keyword(s):

Exact Solution ◽

Bellman Equation ◽

Bandit Problem ◽

Myopic Policy

We consider the symmetric Poissonian two-armed bandit problem. For the case of switching arms, only one of which creates reward, we solve explicitly the Bellman equation for a β-discounted reward and prove that a myopic policy is optimal.

Download Full-text

A doscounted uniform one-armed bandit problem

Sequential Analysis ◽

10.1080/07474949208836241 ◽

1992 ◽

Vol 11 (1) ◽

pp. 1-15 ◽

Cited By ~ 1

Author(s):

Toshio Hamada

Keyword(s):

Bandit Problem

Download Full-text