Sequential resource allocation in a stochastic environment: an overview and numerical experiments-Reference-Cited by-同舟云学术

Sequential resource allocation in a stochastic environment: an overview and numerical experiments

Published:2021 Issue:3 Volume: Page:13-25
ISSN:1812-5409
Container-title:Bulletin of Taras Shevchenko National University of Kyiv. Series: Physics and Mathematics
language:
Short-container-title:BKNUPhM

Author:

Dzhoha A. S.,

Abstract

In this paper, we consider policies for the sequential resource allocation under the multi-armed bandit problem in a stochastic environment. In this model, an agent sequentially selects an action from a given set and an environment reveals a reward in return. In the stochastic setting, each action is associated with a probability distribution with parameters that are not known in advance. The agent makes a decision based on the history of the chosen actions and obtained rewards. The objective is to maximize the total cumulative reward, which is equivalent to the loss minimization. We provide a brief overview of the sequential analysis and an appearance of the multi-armed bandit problem as a formulation in the scope of the sequential resource allocation theory. Multi-armed bandit classification is given with an analysis of the existing policies for the stochastic setting. Two different approaches are shown to tackle the multi-armed bandit problem. In the frequentist view, the confidence interval is used to express the exploration-exploitation trade-off. In the Bayesian approach, the parameter that needs to be estimated is treated as a random variable. Shown, how this model can be modelled with help of the Markov decision process. In the end, we provide numerical experiments in order to study the effectiveness of these policies.

Publisher

Taras Shevchenko National University of Kyiv

Subject

Medical Assisting and Transcription,Medical Terminology

Reference33 articles.

1. 1. WALD, A. (1950) Sequential Analysis. John Wiley & Sons, Inc., NY.

2. On a method of estimating frequencies;HALDANE;Biometrika 33 (3),1945

3. A two-sample test for a li- near hypothesis whose power is independent of the variance;STEIN;The Annals of Mathematical Statistics 16 (3),1945

4. Optimum character of the sequential probability ratio test;WALD;Ann Math Statist,1948

5. Bayes and mini- max solutions of sequential decision problems;ARROW;Econometrica 17,1949