Online Learning for Dual-Index Policies in Dual-Sourcing Systems-Reference-Cited by-同舟云学术

Online Learning for Dual-Index Policies in Dual-Sourcing Systems

Published:2023-12-12 Issue: Volume: Page:
ISSN:1523-4614
Container-title:Manufacturing & Service Operations Management
language:en
Short-container-title:M&SOM

Author:

Tang Jingwen¹^ORCID,Chen Boxiao²^ORCID,Shi Cong³^ORCID

Affiliation:

1. Industrial and Operations Engineering, University of Michigan, Ann Arbor, Michigan 48109;

2. College of Business Administration, University of Illinois Chicago, Chicago, Illinois 60607;

3. Management Science, Miami Herbert Business School, University of Miami, Coral Gables, Florida 33146

Abstract

Problem definition: We consider a periodic-review dual-sourcing inventory system with a regular source (lower unit cost but longer lead time) and an expedited source (shorter lead time but higher unit cost) under carried-over supply and backlogged demand. Unlike existing literature, we assume that the firm does not have access to the demand distribution a priori and relies solely on past demand realizations. Even with complete information on the demand distribution, it is well known in the literature that the optimal inventory replenishment policy is complex and state dependent. Therefore, we focus our attention on a class of popular, easy-to-implement, and near-optimal heuristic policies called the dual-index policy. Methodology/results: The performance measure is the regret, defined as the cost difference of any feasible learning algorithm against the full-information optimal dual-index policy. We develop a nonparametric online learning algorithm that admits a regret upper bound of [Formula: see text], which matches the regret lower bound for any feasible learning algorithms up to a logarithmic factor. Our algorithm integrates stochastic bandits and sample average approximation techniques in an innovative way. As part of our regret analysis, we explicitly prove that the underlying Markov chain is ergodic and converges to its steady state exponentially fast via coupling arguments, which could be of independent interest. Managerial implications: Our work provides practitioners with an easy-to-implement, robust, and provably good online decision support system for managing a dual-sourcing inventory system. Funding: This work was supported by the Amazon Research Award. Supplemental Material: The online appendix is available at https://doi.org/10.1287/msom.2022.0323 .

Publisher

Institute for Operations Research and the Management Sciences (INFORMS)

Subject

Management Science and Operations Research,Strategy and Management

Link

https://pubsonline.informs.org/doi/pdf/10.1287/msom.2022.0323

Reference35 articles.

1. Learning in Structured MDPs with Convex Cost Functions: Improved Regret Bounds for Inventory Management

2. Global Dual Sourcing: Tailored Base-Surge Allocation to Near- and Offshore Production

3. Non-Stationary Stochastic Optimization

4. Some Results Concerning Optimum Inventory Policies

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Safety Stock Policy for Dual Sourcing Inventory System with Forecast Evolution;2024