Scheduling Jobs That Are Subject to Deterministic Due Dates and Have Deteriorating Expected Rewards

A single server processes jobs that can yield rewards but expire on predetermined dates. Expected immediate rewards from each job are deteriorating. The instance is formulated as a multiarmed bandit problem, and an index-based scheduling policy is shown to maximize the expected total reward.

Download Full-text

Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-Part I: I.I.D. rewards

IEEE Transactions on Automatic Control ◽

10.1109/tac.1987.1104491 ◽

1987 ◽

Vol 32 (11) ◽

pp. 968-976 ◽

Cited By ~ 102

Author(s):

V. Anantharam ◽

P. Varaiya ◽

J. Walrand

Keyword(s):

Efficient Allocation ◽

Bandit Problem ◽

Allocation Rules ◽

Asymptotically Efficient ◽

Multiarmed Bandit

Download Full-text

Comments on “Finite-Time Analysis of the Multiarmed Bandit Problem”

2019 International Conference on Machine Learning and Cybernetics (ICMLC) ◽

10.1109/icmlc48188.2019.8949232 ◽

2019 ◽

Cited By ~ 1

Author(s):

Lu-Ning Zhang ◽

Xin Zuo ◽

Jian-Wei Liu ◽

Wei-Min Li ◽

Nobuyasu Ito

Keyword(s):

Finite Time ◽

Bandit Problem ◽

Time Analysis ◽

Multiarmed Bandit

Download Full-text

Index policies for discounted bandit problems with availability constraints

Advances in Applied Probability ◽

10.1017/s0001867800002573 ◽

2008 ◽

Vol 40 (02) ◽

pp. 377-400 ◽

Cited By ~ 1

Author(s):

Savas Dayanik ◽

Warren Powell ◽

Kazutoshi Yamazaki

Keyword(s):

Bandit Problem ◽

Bandit Problems ◽

Index Policy ◽

State Action ◽

Index Policies ◽

Availability Constraints ◽

Whittle Index ◽

Multiarmed Bandit

A multiarmed bandit problem is studied when the arms are not always available. The arms are first assumed to be intermittently available with some state/action-dependent probabilities. It is proven that no index policy can attain the maximum expected total discounted reward in every instance of that problem. The Whittle index policy is derived, and its properties are studied. Then it is assumed that the arms may break down, but repair is an option at some cost, and the new Whittle index policy is derived. Both problems are indexable. The proposed index policies cannot be dominated by any other index policy over all multiarmed bandit problems considered here. Whittle indices are evaluated for Bernoulli arms with unknown success probabilities.

Download Full-text

A Marginal Productivity Index Policy for the Finite-Horizon Multiarmed Bandit Problem

Proceedings of the 44th IEEE Conference on Decision and Control ◽

10.1109/cdc.2005.1582407 ◽

2006 ◽

Cited By ~ 2

Author(s):

J. Nino-Mora

Keyword(s):

Finite Horizon ◽

Productivity Index ◽

Bandit Problem ◽

Marginal Productivity ◽

Index Policy ◽

Multiarmed Bandit

Download Full-text

Asymptotically efficient allocation rules for the multiarmed bandit problem with multiple plays-Part II: Markovian rewards

IEEE Transactions on Automatic Control ◽

10.1109/tac.1987.1104485 ◽

1987 ◽

Vol 32 (11) ◽

pp. 977-982 ◽

Cited By ~ 80

Author(s):

V. Anantharam ◽

P. Varaiya ◽

J. Walrand

Keyword(s):

Efficient Allocation ◽

Bandit Problem ◽

Allocation Rules ◽

Asymptotically Efficient ◽

Multiarmed Bandit

Download Full-text

A FLUID LIMIT FOR PROCESSOR-SHARING QUEUES WEIGHTED BY FUNCTIONS OF REMAINING AMOUNTS OF SERVICE

Probability in the Engineering and Informational Sciences ◽

10.1017/s0269964819000457 ◽

2019 ◽

pp. 1-8

Author(s):

Yingdong Lu

Keyword(s):

Single Server ◽

Processor Sharing ◽

Single Server Queue ◽

Fluid Limit ◽

Scheduling Policy ◽

Server Queue ◽

Processor Sharing Queues

Abstract We study a single server queue under a processor-sharing type of scheduling policy, where the weights for determining the sharing are given by functions of each job's remaining service (processing) amount, and obtain a fluid limit for the scaled measure-valued system descriptors.

Download Full-text

Matching While Learning

Operations Research ◽

10.1287/opre.2020.2013 ◽

2021 ◽

Author(s):

Ramesh Johari ◽

Vijay Kamble ◽

Yash Kanoria

Keyword(s):

Labor Market ◽

Shadow Price ◽

Cold Start ◽

Bandit Problem ◽

Different Types ◽

The Future ◽

Limited Supply ◽

Cold Starts ◽

Cold Start Problem ◽

Multiarmed Bandit

Platforms face a cold start problem whenever new users arrive: namely, the platform must learn attributes of new users (explore) in order to match them better in the future (exploit). How should a platform handle cold starts when there are limited quantities of the items being recommended? For instance, how should a labor market platform match workers to jobs over the lifetime of the worker, given a limited supply of jobs? In this setting, there is one multiarmed bandit problem for each worker, coupled together by the constrained supply of jobs of different types. A solution is developed to this problem. It is found that the platform should estimate a shadow price for each job type, and for each worker, adjust payoffs by these prices (i) to balance learning with payoffs early on and (ii) to myopically match them thereafter.

Download Full-text

A Note on the Equivalence of Upper Confidence Bounds and Gittins Indices for Patient Agents

Operations Research ◽

10.1287/opre.2020.1987 ◽

2020 ◽

Author(s):

Daniel Russo

Keyword(s):

Posterior Distribution ◽

Error Term ◽

Discount Factor ◽

Gittins Index ◽

Confidence Bound ◽

Bandit Problem ◽

Confidence Bounds ◽

Upper Confidence Bound ◽

Gittins Indices ◽

Multiarmed Bandit

This note gives a short, self-contained proof of a sharp connection between Gittins indices and Bayesian upper confidence bound algorithms. I consider a Gaussian multiarmed bandit problem with discount factor [Formula: see text]. The Gittins index of an arm is shown to equal the [Formula: see text]-quantile of the posterior distribution of the arm's mean plus an error term that vanishes as [Formula: see text]. In this sense, for sufficiently patient agents, a Gittins index measures the highest plausible mean-reward of an arm in a manner equivalent to an upper confidence bound.

Download Full-text

An Online Minimax Optimal Algorithm for Adversarial Multiarmed Bandit Problem

IEEE Transactions on Neural Networks and Learning Systems ◽

10.1109/tnnls.2018.2806006 ◽

2018 ◽

Vol 29 (11) ◽

pp. 5565-5580 ◽

Cited By ~ 1

Author(s):

Kaan Gokcesu ◽

Suleyman Serdar Kozat

Keyword(s):

Optimal Algorithm ◽

Bandit Problem ◽

Multiarmed Bandit

Download Full-text

Arm order recognition in multi-armed bandit problem with laser chaos time series

Scientific Reports ◽

10.1038/s41598-021-83726-8 ◽

2021 ◽

Vol 11 (1) ◽

Author(s):

Naoki Narisawa ◽

Nicolas Chauvet ◽

Mikio Hasegawa ◽

Makoto Naruse

Keyword(s):

Time Series ◽

Recognition Accuracy ◽

Information And Communications Technology ◽

Estimation Accuracy ◽

Communications Technology ◽

Allocation Of Resources ◽

Bandit Problem ◽

Scalable Algorithm ◽

Total Reward ◽

Time Division

AbstractBy exploiting ultrafast and irregular time series generated by lasers with delayed feedback, we have previously demonstrated a scalable algorithm to solve multi-armed bandit (MAB) problems utilizing the time-division multiplexing of laser chaos time series. Although the algorithm detects the arm with the highest reward expectation, the correct recognition of the order of arms in terms of reward expectations is not achievable. Here, we present an algorithm where the degree of exploration is adaptively controlled based on confidence intervals that represent the estimation accuracy of reward expectations. We have demonstrated numerically that our approach did improve arm order recognition accuracy significantly, along with reduced dependence on reward environments, and the total reward is almost maintained compared with conventional MAB methods. This study applies to sectors where the order information is critical, such as efficient allocation of resources in information and communications technology.

Download Full-text