Article Open Access

From Batch Prediction to Self-Healing ML Systems: A Production MLOps Framework for Autonomous Model Maintenance

Kushwanth Chowdary Kandala

Abstract


Machine learning (ML) systems deployed in large-scale healthcare Software-as-a-Service (SaaS) environments frequently experience performance degradation after production release due to data drift, evolving data distributions, pipeline latency anomalies, and declining predictive accuracy. These issues often remain undetected until they significantly affect operational performance, resulting in prolonged incident detection, costly engineering interventions, and increased business risk. This study proposes a self-healing machine learning framework that transforms conventional batch-oriented ML pipelines into autonomous, event-driven operational systems capable of continuously monitoring, diagnosing, and recovering from production failures. The proposed architecture integrates four complementary capabilities: statistical data drift detection, automated retraining triggers, canary-based deployment and promotion workflows, and multi-tier observability that connects model behaviour with automated operational responses. Together, these components establish a closed feedback loop that enables continuous adaptation while minimizing manual intervention. The framework is evaluated through a production case study involving a healthcare long-term care SaaS platform supporting large-scale clinical operations. Empirical results demonstrate substantial operational improvements following deployment of the self-healing architecture. Mean time to detect model drift decreased from 18.4 hours to 2.1 hours, while mean time to recovery was reduced from 9.2 hours to 1.8 hours. In addition, the number of monthly incidents requiring manual intervention declined from 34 to 6, indicating significant gains in operational resilience and engineering efficiency. The proposed framework is implemented using widely adopted open-source technologies, including MLflow, Kubeflow Pipelines, Feast, Great Expectations, and Argo Rollouts, allowing seamless integration with existing Kubernetes-based enterprise infrastructures. The findings demonstrate that autonomous, event-driven maintenance substantially improves the reliability, scalability, and maintainability of production ML systems, providing a practical engineering architecture for resilient AI operations in mission-critical healthcare environments

Keywords


MLOps, Self-Healing Systems, Data Drift, Model Monitoring, Healthcare AI

References


B. Eken, S. Pallewatta, N. Tran, A. Tosun, and M. A. Babar, "A multivocal review of MLOps practices, challenges, and open issues," ACM Comput. Surv., vol. 58, no. 2, pp. 1–35, 2025.

S. Amershi et al., "Software engineering for machine learning: A case study," in 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2019, pp. 291–300.

D. Sculley et al., "Hidden technical debt in machine learning systems," Adv. Neural Inf. Process. Syst., vol. 28, 2015.

B. Yurdakul, Statistical properties of population stability index. Western Michigan University, 2018.

M. Mitchell et al., "Model cards for model reporting," in Proceedings of the conference on fairness, accountability, and transparency, 2019, pp. 220–229.

M. R. Miah, "Advanced Computing Frameworks for Distributed Training, Deployment, and Monitoring of Artificial Intelligence and Machine Learning Models," ASRC Procedia Glob. Perspect. Sci. Scholarship, vol. 1, no. 01, pp. 1922–1957, 2025.

Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan, "Dissecting racial bias in an algorithm used to manage the health of populations," Science (80-)., vol. 366, no. 6464, pp. 447–453, 2019.

A. Satyanarayanan, "Foundational Framework Self-Healing Data Pipelines for AI Engineering: A Framework and Implementation," Int. J. Artif. Intell. Data Sci. Mach. Learn., vol. 3, no. 1, pp. 63–76, 2022.

Z. He et al., "Clinical trial generalizability assessment in the big data era: a review," Clin. Transl. Sci., vol. 13, no. 4, pp. 675–684, 2020.

D. Kreuzberger, N. Kühl, and S. Hirschl, "Machine learning operations (mlops): Overview, definition, and architecture," IEEE Access, vol. 11, pp. 31866–31879, 2023.

E. Breck, S. Cai, E. Nielsen, M. Salib, and D. Sculley, "The ML test score: A rubric for ML production readiness and technical debt reduction," in 2017 IEEE International Conference on Big Data (Big Data), 2017, pp. 1123–1132.

G. Hovakimyan and J. M. Bravo, "Evolving strategies in machine learning: a systematic review of concept drift detection," Information, vol. 15, no. 12, p. 786, 2024.

E. A. June, "ML Model Monitoring: How to Detect Data Drift and Model Decay in Production".

J. Lu, A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang, "Learning under concept drift: A review," IEEE Trans. Knowl. Data Eng., vol. 31, no. 12, pp. 2346–2363, 2018.

M. Baena-Garcia, J. del Campo-Ávila, R. Fidalgo, A. Bifet, R. Gavalda, and R. Morales-Bueno, "Early drift detection method," in Fourth international workshop on knowledge discovery from data streams, 2006, pp. 77–86.

V. Losing, B. Hammer, and H. Wersing, "Incremental online learning: A review and comparison of state-of-the-art algorithms," Neurocomputing, vol. 275, pp. 1261–1274, 2018.

M. De Lange et al., "A continual learning survey: Defying forgetting in classification tasks," IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 7, pp. 3366–3385, 2021.

S. Proksch, S. Nadi, S. Amann, and M. Mezini, "Enriching in-IDE process information with fine-grained source code history," in 2017 IEEE 24th International Conference on Software Analysis, Evolution and Reengineering (SANER), 2017, pp. 250–260.

D. Sato, A. Wider, and C. Windheuser, "Continuous delivery for machine learning," Martin Fowler, vol. 9, 2019.

R. Thoppilan et al., "Lamda: Language models for dialog applications," arXiv Prepr. arXiv2201.08239, 2022.

S. Y. Shah et al., "AutoAI-TS: AutoAI for time series forecasting," in Proceedings of the 2021 International Conference on Management of Data, 2021, pp. 2584–2596.

I. C. Vladu, N. G. B^izdoacua, I. Pirici, T.-A. Bual?eanu, and E. N. Bondoc, "The DIME Architecture: A Unified Operational Algorithm for Neural Representation, Dynamics, Control and Integration," Appl. Sci., vol. 16, no. 11, p. 5380, 2026.

L. Zhu, Q. Lu, M. Ding, S. U. Lee, and C. Wang, "Designing meaningful human oversight in AI," AI Ethics, vol. 6, no. 3, p. 286, 2026.

H. M. Wong, L. L. Chan, L. J. Khoo, and S. Perumal, "An End-to-End Adaptive Pipeline for Drift-Aware Model Retraining: Integrating Data Drift Detection and Data Quality Validation," in 2026 IEEE 2nd International Conference on Robotics and Technologies for Industrial Automation (ROBOTHIA), 2026, pp. 1–6.

K. Nguyen, "A Hybrid Markov Decision Process (MDP) Model for Predicting Liquidity Profiles of US Regional Banks," The George Washington University, 2026.

J. Gama, I. Žliobait?, A. Bifet, M. Pechenizkiy, and A. Bouchachia, "A survey on concept drift adaptation," ACM Comput. Surv., vol. 46, no. 4, pp. 1–37, 2014.

N. Polyzotis, M. Zinkevich, S. Roy, E. Breck, and S. Whang, "Data validation for machine learning," Proc. Mach. Learn—Syst., vol. 1, pp. 334–347, 2019.

S. Rukh, O. B. Seyi-Lande, and S. T. Oziri, "Framework design for machine learning adoption in enterprise performance optimization," Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol., vol. 8, no. 3, pp. 798–830, 2022.

G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, "Improving neural networks by preventing co-adaptation of feature detectors," arXiv Prepr. arXiv1207.0580, 2012.

J. M. Zhang, M. Harman, L. Ma, and Y. Liu, "Machine learning testing: Survey, landscapes and horizons," IEEE Trans. Softw. Eng., vol. 48, no. 1, pp. 1–36, 2020.

R. Pandipati, "CLOUD-NATIVE DATA ANALYTICS PLATFORM WITH INTEGRATED GOVERNANCE: A MODERN APPROACH TO REAL-TIME STREAM PROCESSING AND FEATURE ENGINEERING," Int. J. Cloud Comput., 2024.

O. Cobb and A. Van Looveren, "Context-aware drift detection," in International Conference on Machine Learning, 2022, pp. 4087–4111.

C. Zhang et al., "A deep neural network for unsupervised anomaly detection and diagnosis in multivariate time series data," in Proceedings of the AAAI conference on artificial intelligence, 2019, pp. 1409–1416.




DOI: https://doi.org/10.52088/ijesty.v6i3.1869

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Kushwanth Chowdary Kandala

International Journal of Engineering, Science, and Information Technology (IJESTY) eISSN 2775-2674