Distributionally Robust Markov Decision Processes and Their Connection to Risk Measures-Reference-Cited by-同舟云学术

Distributionally Robust Markov Decision Processes and Their Connection to Risk Measures

Published:2021-11-23 Issue: Volume: Page:
ISSN:0364-765X
Container-title:Mathematics of Operations Research
language:en
Short-container-title:Mathematics of OR

Author:

Bäuerle Nicole¹^ORCID,Glauner Alexander¹^ORCID

Affiliation:

1. Department of Mathematics, Karlsruhe Institute of Technology, 76128 Karlsruhe, Germany

Abstract

We consider robust Markov decision processes with Borel state and action spaces, unbounded cost, and finite time horizon. Our formulation leads to a Stackelberg game against nature. Under integrability, continuity, and compactness assumptions, we derive a robust cost iteration for a fixed policy of the decision maker and a value iteration for the robust optimization problem. Moreover, we show the existence of deterministic optimal policies for both players. This is in contrast to classical zero-sum games. In case the state space is the real line, we show under some convexity assumptions that the interchange of supremum and infimum is possible with the help of Sion’s minimax theorem. Further, we consider the problem with special ambiguity sets. In particular, we are able to derive some cases where the robust optimization problem coincides with the minimization of a coherent risk measure. In the final section, we discuss two applications: a robust linear-quadratic problem and a robust problem for managing regenerative energy.

Publisher

Institute for Operations Research and the Management Sciences (INFORMS)

Subject

Management Science and Operations Research,Computer Science Applications,General Mathematics

Reference37 articles.

1. Entropic Value-at-Risk: A New Coherent Risk Measure

2. Optimal Management of Wind Energy with Storage: Structural Implications for Policy and Market Design

3. Convexity and Optimization in Banach Spaces

4. Optimal dividend payout model with risk sensitive preferences

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Robust Q-learning algorithm for Markov decision processes under Wasserstein uncertainty;Automatica;2024-10

2. Markov decision processes with risk-sensitive criteria: an overview;Mathematical Methods of Operations Research;2024-04

3. Dynamic QoS Aware Service Composition Framework Based on AHP and Hierarchical Markov Decision Making;IEEE Access;2024

4. OPTIMAL INVESTMENT UNDER PARTIAL INFORMATION AND ROBUST VAR-TYPE CONSTRAINT;International Journal of Theoretical and Applied Finance;2023-08

5. Reinforcement learning with dynamic convex risk measures;Mathematical Finance;2023-04-17