Noise-Aware Delay-Compensated RL for Robust Thermal Fan Control
DOI: https://doi.org/10.62381/I265609
Author(s)
Shiqi Zhang1, Wen Li1,*, Jingyan Min1, Weibo Tao2, Yazhi Yang1
Affiliation(s)
1School of Electrical Engineering, Yingkou Institute of Technology, Yingkou, China
2School of Mechanical and Power Engineering, Yingkou Institute of Technology, Yingkou, China
*Corresponding Author.
Abstract
Thermal management systems in electronics and data centers rely on closed-loop fan control to maintain safe operating temperatures. In practice, temperature sensors exhibit measurement noise and communication delays that degrade controller performance, yet most existing approaches assume access to clean, instantaneous state observations. We propose a unified framework that integrates Extended Kalman Filter state estimation with delay-compensated deep reinforcement learning to achieve robust thermal regulation under realistic sensor imperfections. The framework employs multi-step state prediction to compensate for measurement latency, while domain randomization during training ensures robustness across diverse noise and delay conditions without retuning. Through comprehensive simulated evaluation across noise levels (σ∈[0.1,5]°C) and measurement delays (0–2s), we demonstrate that the proposed method achieves 47% lower tracking error than PID control and approximately 27% lower energy consumption under challenging conditions (σ=5°C, 400 ms delay), while maintaining computational overhead of only 2.6 times that of PID. The learned policy generalizes to unseen thermal plant dynamics without retraining, suggesting practical applicability to heterogeneous deployment environments.
Keywords
Thermal Management; Fan Control; Sensor Noise; Measurement Delay; Extended Kalman Filter; Deep Reinforcement Learning
References
[1] Tarragona J, de Gracia A, Cabeza LF. Bibliometric analysis of smart control applications in thermal energy storage systems: A model predictive control approach. Journal of Energy Storage, 2020, 32: 101704.
[2] Pratik P, Arulselvan M, Kumar R. Thermal design current control mechanism using PID controller in modern processors. In: 2023 IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT), 2023: 1-5.
[3] Serrano JNP, Martinez AAP, Perez CDZ, et al. Design and implementation of a low-cost thermal chamber with PID control and Kalman filter state estimation. International Journal of Combinatorial Optimization Problems and Informatics, 2025, 16(4): 79-95.
[4] Tan F, Li HX, Shen P. Smith predictor-based multiple periodic disturbance compensation for long dead-time processes. International Journal of Control, 2018, 91(5): 999-1010.
[5] Bai Z, Liu A, Duan Z. Research on the control strategy of electric vehicle battery and cabin thermal management system based on fuzzy PID. In: Ninth International Symposium on Advances in Electrical, Electronics, and Computer Engineering (ISAEECE 2024), SPIE, 2024: 195.
[6] Tarragona J, Pisello AL, Fernandez C, de Gracia A, Cabeza LF. Systematic review on model predictive control strategies applied to active thermal energy storage systems. Renewable and Sustainable Energy Reviews, 2021, 149: 111385.
[7] Lu R, Li X, Chen R, Lei A, Ma X. An alternative reinforcement learning (ARL) control strategy for data center air-cooled HVAC systems. Energy, 2024, 308: 132977.
[8] Yang W, Xu Y. A deep reinforcement learning framework for optimizing data center cooling systems. In: 2025 10th International Conference on Power and Renewable Energy (ICPRE), IEEE, 2025: 1796-1801.
[9] Mu N, Hu X, Jia QS. Integrating mechanism and data: Reinforcement learning based on multi-fidelity model for data center cooling control. In: 2023 China Automation Congress (CAC), IEEE, 2023: 5283-5288.
[10] Chen X, Hu J, Jin C, Li L, Wang L. Understanding domain randomization for sim-to-real transfer. In: The Tenth International Conference on Learning Representations (ICLR), 2022.
[11] Tiboni G, Arndt K, Kyrki V. DROPO: Sim-to-real transfer with offline domain randomization. Robotics and Autonomous Systems, 2023, 166: 104432.
[12] Altaf M, Elias M, Feng Y. Model predictive control for active thermal management of autonomous electric vehicles. IEEE Transactions on Control Systems Technology, 2017, 26(6): 2045-2057.
[13] Tang M, Cai S, Lau VKN. Temperature control for cyber-physical thermal systems over wireless networks: A model-assisted deep reinforcement learning approach. IEEE Transactions on Signal Processing, 2026, 74: 1-15.
[14] Ding L, Wen C. High-order extended Kalman filter for state estimation of nonlinear systems. Symmetry, 2024, 16(5): 617.
[15] Murakami R, Mori S, Zhang HK. Thermal ablation therapy control with tissue necrosis-driven temperature feedback enabled by neural state space model with extended Kalman filter. In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024: 2373-2379.
[16] Nobahari H, Sharifi A. A hybridization of extended Kalman filter and ant colony optimization for state estimation of nonlinear systems. Applied Soft Computing, 2019, 74: 411-423.
[17] Smith OJM. A controller to overcome dead time. ISA Journal, 1957, 6(2): 28-33.
[18] Martinovic M, Varol HA. Delay-aware model predictive control for thermal systems. IEEE Transactions on Control Systems Technology, 2020, 28(3): 1082-1089.
[19] Walsh TJ, Nouri A, Li L, Littman ML. Learning and planning with timing information in Markov decision processes. In: Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI), 2009: 555-562.
[20] Haarnoja T, Zhou A, Abbeel P, Levine S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: Proceedings of the 35th International Conference on Machine Learning (ICML), 2018: 1861-1870.