Convergence of a Q-learning Variant for Continuous States and Actions-Reference-Cited by-同舟云学术

Convergence of a Q-learning Variant for Continuous States and Actions

Published:2014-04-29 Issue: Volume:49 Page:705-731
ISSN:1076-9757
Container-title:Journal of Artificial Intelligence Research
language:
Short-container-title:jair

Author:

Carden S. W.

Abstract

This paper presents a reinforcement learning algorithm for solving infinite horizon Markov Decision Processes under the expected total discounted reward criterion when both the state and action spaces are continuous. This algorithm is based on Watkins' Q-learning, but uses Nadaraya-Watson kernel smoothing to generalize knowledge to unvisited states. As expected, continuity conditions must be imposed on the mean rewards and transition probabilities. Using results from kernel regression theory, this algorithm is proven capable of producing a Q-value function estimate that is uniformly within an arbitrary tolerance of the true Q-value function with probability one. The algorithm is then applied to an example problem to empirically show convergence as well.

Publisher

AI Access Foundation

Subject

Artificial Intelligence

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Using Reinforcement Learning to Minimize the Probability of Delay Occurrence in Transportation;IEEE Transactions on Vehicular Technology;2020-03

2. Exploration Using Without-Replacement Sampling of Actions Is Sometimes Inferior;Machine Learning and Knowledge Extraction;2019-05-24

3. Observation-Based Optimization for POMDPs With Continuous State, Observation, and Action Spaces;IEEE Transactions on Automatic Control;2019-05

4. Small-sample reinforcement learning: Improving policies using synthetic data1;Intelligent Decision Technologies;2017-06-22

5. The Challenges of Reinforcement Learning in Robotics and Optimal Control;Advances in Intelligent Systems and Computing;2016-10-18