Expected scalarised returns dominance: a new solution concept for multi-objective decision making-Reference-Cited by-同舟云学术

Expected scalarised returns dominance: a new solution concept for multi-objective decision making

Published:2022-07-05 Issue: Volume: Page:
ISSN:0941-0643
Container-title:Neural Computing and Applications
language:en
Short-container-title:Neural Comput & Applic

Author:

Hayes Conor F.,Verstraeten Timothy,Roijers Diederik M.,Howley Enda,Mannion Patrick

Abstract

AbstractIn many real-world scenarios, the utility of a user is derived from a single execution of a policy. In this case, to apply multi-objective reinforcement learning, the expected utility of the returns must be optimised. Various scenarios exist where a user’s preferences over objectives (also known as the utility function) are unknown or difficult to specify. In such scenarios, a set of optimal policies must be learned. However, settings where the expected utility must be maximised have been largely overlooked by the multi-objective reinforcement learning community and, as a consequence, a set of optimal solutions has yet to be defined. In this work, we propose first-order stochastic dominance as a criterion to build solution sets to maximise expected utility. We also define a new dominance criterion, known as expected scalarised returns (ESR) dominance, that extends first-order stochastic dominance to allow a set of optimal policies to be learned in practice. Additionally, we define a new solution concept called the ESR set, which is a set of policies that are ESR dominant. Finally, we present a new multi-objective tabular distributional reinforcement learning (MOTDRL) algorithm to learn the ESR set in multi-objective multi-armed bandit settings.

Funder

National University Ireland, Galway

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence,Software

Link

https://link.springer.com/content/pdf/10.1007/s00521-022-07334-x.pdf

Reference49 articles.

1. Ali MM (1975) Stochastic dominance and portfolio analysis. J Finan Econ 2(2): 205–229. https://doi.org/10.1016/0304-405X(75)90005-7. https://www.sciencedirect.com/science/article/pii/0304405X75900057

2. Atkinson AB, Bourguignon F (1982) The comparison of multi-dimensioned distributions of economic status. Rev Econ Stud 49(2):183–201. https://doi.org/10.2307/2297269

3. Auer P, Chiang CK, Ortner R, Drugan M (2016) Pareto front identification from stochastic bandit feedback. In: Gretton A, Robert CC (eds) Proceedings of the 19th international conference on artificial intelligence and statistics, proceedings of machine learning research, vol 51, pp 939–947. PMLR, Cadiz, Spain. http://proceedings.mlr.press/v51/auer16.html

4. Bawa VS (1975) Optimal rules for ordering uncertain prospects. J Finan Econ 2(1): 95–121. https://doi.org/10.1016/0304-405X(75)90025-2. http://www.sciencedirect.com/science/article/pii/0304405X75900252

5. Bawa VS (1978) Safety-first, stochastic dominance, and optimal portfolio choice. J Finan Quant Anal 13(2): 255–271. http://www.jstor.org/stable/2330386

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Multi-objective intelligent clustering routing schema for internet of things enabled wireless sensor networks using deep reinforcement learning;Cluster Computing;2024-01-09

2. Actor-critic multi-objective reinforcement learning for non-linear utility functions;Autonomous Agents and Multi-Agent Systems;2023-04-28

3. Monte Carlo tree search algorithms for risk-aware and multi-objective reinforcement learning;Autonomous Agents and Multi-Agent Systems;2023-04-28

4. Evaluation and optimization of ecological compensation fairness in prefecture-level cities of Anhui province;Environmental Research Communications;2023-03-01