A linear programming based approach for composite-action Markov decision processes-Reference-Cited by-同舟云学术

A linear programming based approach for composite-action Markov decision processes

Published:2019-10-09 Issue:5 Volume:53 Page:1749-1761
ISSN:0399-0559
Container-title:RAIRO - Operations Research
language:
Short-container-title:RAIRO-Oper. Res.

Author:

Zhang Zhicong,Li Shuai,Yan Xiaohui,Zhang Liangwei

Abstract

We study a time homogeneous discrete composite-action Markov decision process (CMDP) which needs to make multiple decisions at each state. In this particular Markov decision process, the state variables are divided into two separable sets and a two-dimensional composite action is chosen at each decision epoch. To solve a composite-action Markov decision process, we propose a novel linear programming model (Contracted Linear Programming Model, CLPM). We show that the CLPM model obtains the optimal state values of a CMDP process. We analyze and compare the number of variables and constraints of the CLPM model and the Traditional Linear Programming Model (TLPM). Computational experiments compare running times and memory usage of the two models. The CLPM model outperforms the TLPM model in both time complexity and space complexity by theoretical analysis and computational experiments.

Funder

Science and Technology Planning Project of Guangdong Province, China

Natural Science Foundation of Guangdong Province, China

National Natural Science Foundation of China

Publisher

EDP Sciences

Subject

Management Science and Operations Research,Computer Science Applications,Theoretical Computer Science

Link

https://www.rairo-ro.org/10.1051/ro/2018081/pdf

Reference14 articles.

1. OPTIMAL CONTROL OF A TWO-STAGE TANDEM QUEUING SYSTEM WITH FLEXIBLE SERVERS

2. Multi-Actor Markov Decision Processes

3. Throughput Maximization for Tandem Lines with Two Stations and Flexible Servers

4. Throughput maximization for two station tandem systems: a proof of the Andradóttir–Ayhan conjecture

5. On the Introduction of an Agile, Temporary Workforce into a Tandem Queueing System