Refined Continuous Control of DDPG Actors via Parametrised Activation-Reference-Cited by-同舟云学术

Refined Continuous Control of DDPG Actors via Parametrised Activation

Published:2021-09-29 Issue:4 Volume:2 Page:464-476
ISSN:2673-2688
Container-title:AI
language:en
Short-container-title:AI

Author:

Hossny Mohammed^ORCID,Iskander Julie^ORCID,Attia Mohamed^ORCID,Saleh Khaled^ORCID,Abobakr Ahmed^ORCID

Abstract

Continuous action spaces impose a serious challenge for reinforcement learning agents. While several off-policy reinforcement learning algorithms provide a universal solution to continuous control problems, the real challenge lies in the fact that different actuators feature different response functions due to wear and tear (in mechanical systems) and fatigue (in biomechanical systems). In this paper, we propose enhancing the actor-critic reinforcement learning agents by parameterising the final layer in the actor network. This layer produces the actions to accommodate the behaviour discrepancy of different actuators under different load conditions during interaction with the environment. To achieve this, the actor is trained to learn the tuning parameter controlling the activation layer (e.g., Tanh and Sigmoid). The learned parameters are then used to create tailored activation functions for each actuator. We ran experiments on three OpenAI Gym environments, i.e., Pendulum-v0, LunarLanderContinuous-v2, and BipedalWalker-v2. Results showed an average of 23.15% and 33.80% increase in total episode reward of the LunarLanderContinuous-v2 and BipedalWalker-v2 environments, respectively. There was no apparent improvement in Pendulum-v0 environment but the proposed method produces a more stable actuation signal compared to the state-of-the-art method. The proposed method allows the reinforcement learning actor to produce more robust actions that accommodate the discrepancy in the actuators’ response functions. This is particularly useful for real life scenarios where actuators exhibit different response functions depending on the load and the interaction with the environment. This also simplifies the transfer learning problem by fine-tuning the parameterised activation layers instead of retraining the entire policy every time an actuator is replaced. Finally, the proposed method would allow better accommodation to biological actuators (e.g., muscles) in biomechanical systems.

Publisher

MDPI AG

Link

https://www.mdpi.com/2673-2688/2/4/29/pdf

Reference38 articles.

1. Artificial Intelligence for Prosthetics: Challenge Solutions;Kidziński,2020

2. Learning to run challenge: Synthesizing physiologically accurate motion using deep reinforcement learning;Kidziński,2018

3. Human-level control through deep reinforcement learning

4. Reinforcement learning in robotics: A survey

5. Adjustment of Muscle Mechanics Model Parameters to Simulate Dynamic Contractions in Older Adults

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Route recommendation method for frequent passengers in subway based on passenger preference ranking;Expert Systems with Applications;2024-10

2. A DRL Strategy for Optimal Resource Allocation Along With 3D Trajectory Dynamics in UAV-MEC Network;IEEE Access;2023

3. Route Recommendation Method for Frequent Passengers in Subway Based on Passenger Preference Ranking;2023