Adaptive Inventory Control under Non-Stationary Retail Demand: a Deep Reinforcement Learning Study on the M5 Dataset
DOI: https://doi.org/10.62381/ACS.CESS2026.08
Author(s)
Yongjia Xie
Affiliation(s)
Xi'an Jiaotong-Liverpool University, Suzhou, China
Abstract
Retail demand is not stable over time. Its average level and its variability change because of trend, seasonality, promotions, and shifts in consumer behavior. A fixed inventory rule that is tuned on past demand can become wrong when the demand level moves to a new range. This paper studies that problem with real data. We take one high volume product from the public M5 Walmart dataset and build a daily, single product, lost sales inventory environment that is driven by the real demand of that product. We train a deep reinforcement learning agent with Proximal Policy Optimization to set the daily order quantity. We compare the learned policy against two classic policies, a base stock policy and an (s, S) policy, both tuned on the training period. The held out test period for this product has a much lower mean demand than the training period, which makes it a clear test of adaptation. The deep reinforcement learning policy lowers the total cost on the test period by 13.5 percent, from 19705 to 17052 in the same cost units. Almost all of this gain comes from holding cost, which the agent cuts by about 48 percent because it stops over stocking once demand falls. The cost gain comes with a lower service level. The fill rate of the learned policy is 88.2 percent against 99.1 percent for the fixed rule, which reflects the chosen cost weights. The result shows that an adaptive learned policy can track a new demand level and reduce cost, while a fixed policy keeps acting on an old level. We report the full setup, the data statistics, and the results so the study can be reproduced.
Keywords
Inventory Optimization; Non-Stationary Demand; Deep Reinforcement Learning; Proximal Policy Optimization; M5 Dataset; Retail Supply Chain; Service Level
References
[1] Dehaybe, H.; Catanzaro, D.; Chevalier, P. Deep Reinforcement Learning for inventory optimization with non-stationary uncertain demand. European Journal of Operational Research, 2024, 314(2), 433-445.
[2] Yang, Y.; Wang, M.; Wang, J.; Li, P.; Zhou, M. Multi-Agent Deep Reinforcement Learning for Integrated Demand Forecasting and Inventory Optimization in Sensor-Enabled Retail Supply Chains. Sensors, 2025, 25(8), 2428. https://doi.org/10.3390/s25082428
[3] Lu, X.; Wang, H.; Peng, Z.; Liao, C.; Liu, C. Dynamic Optimization of Multi-Echelon Supply Chain Inventory Policies Under Disruptive Scenarios: A Deep Reinforcement Learning Approach. Symmetry, 2025, 17(12), 2078. https://doi.org/10.3390/sym17122078
[4] Oroojlooyjadid, A.; Nazari, M.; Snyder, L. V.; Takac, M. A Deep Q-Network for the Beer Game: Deep Reinforcement Learning for Inventory Optimization. Manufacturing and Service Operations Management, 2022, 24(1), 285-304.
[5] Geevers, K.; van Hezewijk, L.; Mes, M. R. K. Multi-echelon inventory optimization using deep reinforcement learning. Central European Journal of Operations Research, 2024, 32, 653-683.
[6] Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347, 2017.
[7] Lim, B.; Arik, S. O.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. International Journal of Forecasting, 2021, 37(4), 1748-1764.
[8] Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Computation, 1997, 9(8), 1735-1780.
[9] Makridakis, S.; Spiliotis, E.; Assimakopoulos, V. The M5 competition: Background, organization, and implementation. International Journal of Forecasting, 2022, 38(4), 1325-1336.
[10] Snyder, L. V.; Atan, Z.; Peng, P.; Rong, Y.; Schmitt, A. J.; Sinsoysal, B. OR/MS models for supply chain disruptions: A review. IIE Transactions, 2016, 48(2), 89-109.
[11] Douaioui, K.; Oucheikh, R.; Benmoussa, O.; Mabrouki, C. Machine Learning and Deep Learning Models for Demand Forecasting in Supply Chain Management: A Critical Review. Applied System Innovation, 2024, 7(5), 93.
[12] Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; Dormann, N. Stable-Baselines3: Reliable Reinforcement Learning Implementations. Journal of Machine Learning Research, 2021, 22(268), 1-8.