Bi-level deep reinforcement learning for PEV decision-making guidance by coordinating transportation-electrification coupled systems-Reference-Cited by-同舟云学术

Bi-level deep reinforcement learning for PEV decision-making guidance by coordinating transportation-electrification coupled systems

Published:2023-01-11 Issue: Volume:10 Page:
ISSN:2296-598X
Container-title:Frontiers in Energy Research
language:
Short-container-title:Front. Energy Res.

Author:

Xing Qiang,Chen Zhong,Wang Ruisheng,Zhang Ziqi

Abstract

The random charging and dynamic traveling behaviors of massive plug-in electric vehicles (PEVs) pose challenges to the efficient and safe operation of transportation-electrification coupled systems (TECSs). To realize real-time scheduling of urban PEV fleet charging demand, this paper proposes a PEV decision-making guidance (PEVDG) strategy based on the bi-level deep reinforcement learning, achieving the reduction of user charging costs while ensuring the stable operation of distribution networks (DNs). For the discrete time-series characteristics and the heterogeneity of decision actions, the FEVDG problem is duly decoupled into a bi-level finite Markov decision process, in which the upper-lower layers are used respectively for charging station (CS) recommendation and path navigation. Specifically, the upper-layer agent realizes the mapping relationship between the environment state and the optimal CS by perceiving the PEV charging requirements, CS equipment resources and DN operation conditions. And the action decision output of the upper-layer is embedded into the state space of the lower-layer agent. Meanwhile, the lower-level agent determines the optimal road segment for path navigation by capturing the real-time PEV state and the transportation network information. Further, two elaborate reward mechanisms are developed to motivate and penalize the decision-making learning of the dual agents. Then two extension mechanisms (i.e., dynamic adjustment of learning rates and adaptive selection of neural network units) are embedded into the Rainbow algorithm based on the DQN architecture, constructing a modified Rainbow algorithm as the solution to the concerned bi-level decision-making problem. The average rewards for the upper-lower levels are ¥ -90.64 and ¥ 13.24 respectively. The average equilibrium degree of the charging service and average charging cost are 0.96 and ¥ 42.45, respectively. Case studies are conducted within a practical urban zone with the TECS. Extensive experimental results show that the proposed methodology improves the generalization and learning ability of dual agents, and facilitates the collaborative operation of traffic and electrical networks.

Publisher

Frontiers Media SA

Subject

Economics and Econometrics,Energy Engineering and Power Technology,Fuel Technology,Renewable Energy, Sustainability and the Environment

Reference34 articles.

1. Dynamic energy scheduling and routing of multiple electric vehicles using deep reinforcement learning;Alqahtani;Energy,2022

2. Optimal electric vehicle charging strategy with Markov decision process and reinforcement learning technique;Ding;IEEE Trans. Ind. Appl.,2020

3. Deep-Reinforcement-Learning-Based autonomous voltage control for power grid operations;Duan;IEEE Trans. Power Syst.,2020

4. Modeling charging behavior of battery electric vehicle drivers: A cumulative prospect theory based approach;Hu;Transp. Res. Part C Emerg. Technol.,2019