Optimal Exploitation of Clustering and History Information in Multi-armed Bandit-Reference-Cited by-同舟云学术

Optimal Exploitation of Clustering and History Information in Multi-armed Bandit

Published:2019-08 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Bouneffouf Djallel¹,Parthasarathy Srinivasan¹,Samulowitz Horst¹,Wistuba Martin¹

Affiliation:

1. IBM Research, Yorktown Heights, NY, USA

Abstract

We consider the stochastic multi-armed bandit problem and the contextual bandit problem with historical observations and pre-clustered arms. The historical observations can contain any number of instances for each arm, and the pre-clustering information is a fixed clustering of arms provided as part of the input. We develop a variety of algorithms which incorporate this offline information effectively during the online exploration phase and derive their regret bounds. In particular, we develop the META algorithm which effectively hedges between two other algorithms: one which uses both historical observations and clustering, and another which uses only the historical observations. The former outperforms the latter when the clustering quality is good, and vice-versa. Extensive experiments on synthetic and real world datasets on Warafin drug dosage and web server selectionfor latency minimization validate our theoretical insights and demonstrate that META is a robust strategy for optimally exploiting the pre-clustering information.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Multi-armed bandits with dependent arms;Machine Learning;2023-11-20

2. A Recommender for the Management of Chronic Pain in Patients Undergoing Spinal Cord Stimulation;2023 IEEE International Conference on Digital Health (ICDH);2023-07

3. Cutting to the chase with warm-start contextual bandits;Knowledge and Information Systems;2023-04-11

4. Hierarchical Unimodal Bandits;Machine Learning and Knowledge Discovery in Databases;2023

5. Improving the Size and Quality of MAP-Elites Containers via Multiple Emitters and Decoders for Urban Logistics;Applications of Evolutionary Computation;2023