Inverse reinforcement learning in contextual MDPs-Reference-Cited by-同舟云学术

Inverse reinforcement learning in contextual MDPs

Published:2021-05-12 Issue:9 Volume:110 Page:2295-2334
ISSN:0885-6125
Container-title:Machine Learning
language:en
Short-container-title:Mach Learn

Author:

Belogolovsky Stav^ORCID,Korsunsky Philip,Mannor Shie,Tessler Chen,Zahavy Tom

Abstract

AbstractWe consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are sampled from a distribution. In addition, although the reward is a function of the context, it is not provided to the agent. Instead, the agent observes demonstrations from an optimal policy. The goal is to learn the reward mapping, such that the agent will act optimally even when encountering previously unseen contexts, also known as zero-shot transfer. We formulate this problem as a non-differential convex optimization problem and propose a novel algorithm to compute its subgradients. Based on this scheme, we analyze several methods both theoretically, where we compare the sample complexity and scalability, and empirically. Most importantly, we show both theoretically and empirically that our algorithms perform zero-shot transfer (generalize to new and unseen contexts). Specifically, we present empirical experiments in a dynamic treatment regime, where the goal is to learn a reward function which explains the behavior of expert physicians based on recorded data of them treating patients diagnosed with sepsis.

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence,Software

Link

https://link.springer.com/content/pdf/10.1007/s10994-021-05984-x.pdf

Reference47 articles.

1. Abbeel, P., & Ng, A.Y.(2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning (pp. 1). ACM.

2. Abbeel, P., & Ng, A.Y. (2005). Exploration and apprenticeship learning in reinforcement learning. In Proceedings of the 22nd International Conference on Machine Learning, ICML ’05 (pp. 1-8). New York, NY, USA: Association for Computing Machinery. ISBN 1595931805. https://doi.org/10.1145/1102351.1102352.

3. Amin, K., Jiang, N., & Singh, S. (2017). Repeated inverse reinforcement learning. Advances in Neural Information Processing Systems, 1815–1824.

4. Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., & Silver, D. (2017). Successor features for transfer in reinforcement learning. Advances in neural information processing systems, 4055–4065.

5. Beck, A., & Teboulle, M. (2003). Mirror descent and nonlinear projected subgradient methods for convex optimization. Operations Research Letters, 31, 167–175.

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Can Machine Learning Personalize Cardiovascular Therapy in Sepsis?;Critical Care Explorations;2024-05

2. Trajectory modeling via random utility inverse reinforcement learning;Information Sciences;2024-03

3. Inverse Reinforcement Learning as the Algorithmic Basis for Theory of Mind: Current Methods and Open Problems;Algorithms;2023-01-19

4. Deep Reinforcement Learning With Explicit Context Representation;IEEE Transactions on Neural Networks and Learning Systems;2023

5. Cross Apprenticeship Learning Framework: Properties and Solution Approaches;IEEE Open Journal of Control Systems;2023