Assimilating Human Feedback from Autonomous Vehicle Interaction in Reinforcement Learning Models

Author:

Fox Richard1,Ludvig Elliot A.1

Affiliation:

1. University of Warwick

Abstract

Abstract A significant challenge for real-world automated vehicles (AVs) is their interaction with human pedestrians. This paper develops a methodology to directly elicit the AV behaviour pedestrians find suitable by collecting quantitative data that can be used to measure and improve an algorithm's performance. Starting with a Deep Q Network (DQN) trained on a simple Pygame/Python-based pedestrian crossing environment, the reward structure was adapted to allow adjustment by human feedback. Feedback was collected by eliciting behavioural judgements collected from people in a controlled environment. The reward was shaped by the inter-action vector, decomposed into feature aspects for relevant behaviours, thereby facilitating both implicit preference selection and explicit task discovery in tandem. Using computational RL and behavioural-science techniques, we harness a formal iterative feedback loop where the rewards are repeatedly adapted based on human behavioural judgments. Experiments were conducted with 124 participants that showed strong initial improvement in the judgement of AV behaviours with the adaptive reward structure. The results indicate that the primary avenue for enhancing vehicle behaviour lies in the predictability of its movements when introduced. More broadly, recognising AV behaviours that receive favourable human judgments can pave the way for enhanced performance.

Publisher

Research Square Platform LLC

Reference19 articles.

1. The social dilemma of autonomous vehicles;Bonnefon JF;Science

2. Pal, A., Philion, J., Liao, Y. H., & Fidler, S. (2020). ‘Emergent Road Rules In Multi-Agent Driving Environments’, ArXiv201110753 Cs, Nov. Accessed: Feb. 23, 2021. [Online]. Available: http://arxiv.org/abs/2011.10753.

3. Negotiating the Traffic: Can Cognitive Science Help Make Autonomous Vehicles a Reality?;Chater N;Trends In Cognitive Sciences

4. Ritchie, O. T. (Oct. 2019). ‘How should autonomous vehicles overtake other drivers?’, Transp. Res. Part F Traffic Psychol. Behav., vol. 66, pp. 406–418, 10.1016/j.trf.2019.09.016.

5. Knox, W. B., Allievi, A., Banzhaf, H., Schmitt, F., & Stone, P. (2021). ‘Reward (Mis)design for Autonomous Driving’, ArXiv210413906 Cs, Apr. Accessed: Jul. 26, 2021. [Online]. Available: http://arxiv.org/abs/2104.13906.

同舟云学术

1.学者识别学者识别

2.学术分析学术分析

3.人才评估人才评估

"同舟云学术"是以全球学者为主线,采集、加工和组织学术论文而形成的新型学术文献查询和分析系统,可以对全球学者进行文献检索和人才价值评估。用户可以通过关注某些学科领域的顶尖人物而持续追踪该领域的学科进展和研究前沿。经过近期的数据扩容,当前同舟云学术共收录了国内外主流学术期刊6万余种,收集的期刊论文及会议论文总量共计约1.5亿篇,并以每天添加12000余篇中外论文的速度递增。我们也可以为用户提供个性化、定制化的学者数据。欢迎来电咨询!咨询电话:010-8811{复制后删除}0370

www.globalauthorid.com

TOP

Copyright © 2019-2024 北京同舟云网络信息技术有限公司
京公网安备11010802033243号  京ICP备18003416号-3