Using Natural Language for Reward Shaping in Reinforcement Learning-Reference-Cited by-同舟云学术

Using Natural Language for Reward Shaping in Reinforcement Learning

Published:2019-08 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Goyal Prasoon¹,Niekum Scott¹,Mooney Raymond J.¹

Affiliation:

1. The University of Texas at Austin

Abstract

Recent reinforcement learning (RL) approaches have shown strong performance in complex domains, such as Atari games, but are highly sample inefficient. A common approach to reduce interaction time with the environment is to use reward shaping, which involves carefully designing reward functions that provide the agent intermediate rewards for progress towards the goal. Designing such rewards remains a challenge, though. In this work, we use natural language instructions to perform reward shaping. We propose a framework that maps free-form natural language instructions to intermediate rewards, that can seamlessly be integrated into any standard reinforcement learning algorithm. We experiment with Montezuma's Revenge from the Atari video games domain, a popular benchmark in RL. Our experiments on a diverse set of 15 tasks demonstrate that for the same number of interactions with the environment, using language-based rewards can successfully complete the task 60% more often, averaged across all tasks, compared to learning without language.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 17 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Transformable Gaussian Reward Function for Socially Aware Navigation Using Deep Reinforcement Learning;Sensors;2024-07-13

2. Absorb What You Need: Accelerating Exploration via Valuable Knowledge Extraction;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

3. Ask-AC: An Initiative Advisor-in-the-Loop Actor–Critic Framework;IEEE Transactions on Systems, Man, and Cybernetics: Systems;2023-12

4. “Do this instead” – Robots that Adequately Respond to Corrected Instructions;ACM Transactions on Human-Robot Interaction;2023-09-22

5. QWriter System for Robot-Assisted Alphabet Acquisition;2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN);2023-08-28