MURM: Utilization of Multi-Views for Goal-Conditioned Reinforcement Learning in Robotic Manipulation-Reference-Cited by-同舟云学术

MURM: Utilization of Multi-Views for Goal-Conditioned Reinforcement Learning in Robotic Manipulation

Published:2023-08-19 Issue:4 Volume:12 Page:119
ISSN:2218-6581
Container-title:Robotics
language:en
Short-container-title:Robotics

Author:

Jang Seongwon¹^ORCID,Jeong Hyemi¹^ORCID,Yang Hyunseok¹

Affiliation:

1. Mechanical Engineering Department, Yonsei University, Seoul 03722, Republic of Korea

Abstract

We present a novel framework, multi-view unified reinforcement learning for robotic manipulation (MURM), which efficiently utilizes multiple camera views to train a goal-conditioned policy for a robot to perform complex tasks. The MURM framework consists of three main phases: (i) demo collection from an expert, (ii) representation learning, and (iii) offline reinforcement learning. In the demo collection phase, we design a scripted expert policy that uses privileged information, such as Cartesian coordinates of a target and goal, to solve the tasks. We add noise to the expert policy to provide sufficient interactive information about the environment, as well as suboptimal behavioral trajectories. We designed three tasks in a Pybullet simulation environment, including placing an object in a desired goal position and picking up various objects that are randomly positioned in the environment. In the representation learning phase, we use a vector-quantized variational autoencoder (VQVAE) to learn a more structured latent representation that makes it feasible to train for RL compared to high-dimensional raw images. We train VQVAE models for each distinct camera view and define the best viewpoint settings for training. In the offline reinforcement learning phase, we use the Implicit Q-learning (IQL) algorithm as our baseline and introduce a separated Q-functions method and dropout method that can be implemented in multi-view settings to train the goal-conditioned policy with supervised goal images. We conduct experiments in simulation and show that the single-view baseline fails to solve complex tasks, whereas MURM is successful.

Publisher

MDPI AG

Subject

Artificial Intelligence,Control and Optimization,Mechanical Engineering

Link

https://www.mdpi.com/2218-6581/12/4/119/pdf

Reference34 articles.

1. Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, MIT Press.

2. Mastering the game of go without human knowledge;Silver;Nature,2017

3. Deep reinforcement learning based moving object grasping;Chen;Inf. Sci.,2021

4. Su, H., Hu, Y., Li, Z., Knoll, A., Ferrigno, G., and De Momi, E. (August, January 31). Reinforcement learning based manipulation skill transferring for robot-assisted minimally invasive surgery. Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France.

5. Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., and Welinder, P. (2018). Multi-goal reinforcement learning: Challenging robotics environments and request for research. arXiv.