Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis-Reference-Cited by-同舟云学术

Thompson Sampling for Stochastic Bandits with Noisy Contexts: An Information-Theoretic Regret Analysis

Published:2024-07-17 Issue:7 Volume:26 Page:606
ISSN:1099-4300
Container-title:Entropy
language:en
Short-container-title:Entropy

Author:

Jose Sharu Theresa¹^ORCID,Moothedath Shana²^ORCID

Affiliation:

1. School of Computer Science, University of Birmingham, Birmingham B15 2TT, UK

2. Department of Electrical Engineering, Iowa State University, Ames, IA 50011, USA

Abstract

We study stochastic linear contextual bandits (CB) where the agent observes a noisy version of the true context through a noise channel with unknown channel parameters. Our objective is to design an action policy that can “approximate” that of a Bayesian oracle that has access to the reward model and the noise channel parameter. We introduce a modified Thompson sampling algorithm and analyze its Bayesian cumulative regret with respect to the oracle action policy via information-theoretic tools. For Gaussian bandits with Gaussian context noise, our information-theoretic analysis shows that under certain conditions on the prior variance, the Bayesian cumulative regret scales as O˜(mT), where m is the dimension of the feature vector and T is the time horizon. We also consider the problem setting where the agent observes the true context with some delay after receiving the reward, and show that delayed true contexts lead to lower regret. Finally, we empirically demonstrate the performance of the proposed algorithms against baselines.

Publisher

MDPI AG

Link

https://www.mdpi.com/1099-4300/26/7/606/pdf

Reference25 articles.

1. Srivastava, V., Reverdy, P., and Leonard, N.E. (2014, January 15–17). Surveillance in an abruptly changing world via multiarmed bandits. Proceedings of the IEEE Conference on Decision and Control (CDC), Los Angeles, CA, USA.

2. On multi-armed bandit designs for dose-finding clinical trials;Aziz;J. Mach. Learn. Res.,2021

3. Distributed algorithms for learning and cognitive medium access with logarithmic regret;Anandkumar;IEEE J. Sel. Areas Commun.,2011

4. Srivastava, V., Reverdy, P., and Leonard, N.E. (2013, January 2–4). On optimal foraging and multi-armed bandits. Proceedings of the Annual Allerton Conference on Communication, Control, and Computing (Allerton), Monticello, IL, USA.

5. Bubeck, S., and Cesa-Bianchi, N. (2012). Regret analysis of stochastic and nonstochastic multi-armed bandit problems. arXiv.