A Unified Observability Framework for Cloud-Native Machine Learning Systems: Architectural Primitives, Lineage Correlation, and Cross-Stage Reliability Objectives
Abstract
The observability gap in modern cloud-native organizations reflects a fundamental disconnect between data engineering teams, which focus on pipeline throughput and data freshness, and machine learning teams, which monitor inference latency and prediction drift. This separation is largely caused by telemetry systems that lack a unified identity model capable of linking datasets, features, model versions, and production environments across the machine learning lifecycle. To address this challenge, this paper proposes a unified observability framework that integrates telemetry primitives, lineage-aware correlation, and end-to-end reliability objectives connecting data quality and freshness with model performance outcomes. The framework introduces a minimal set of universal telemetry tags, including environment, workload identifier, dataset/feature/model version, and execution run identifier, enabling consistent cross-lifecycle correlation and incident analysis. A comparative evaluation is conducted against existing observability solutions, including MLflow integrated with OpenTelemetry, Monte Carlo data observability, and WhyLogs. The results indicate that the proposed framework offers superior capabilities for cross-stage incident attribution by linking failures occurring across data pipelines, feature engineering processes, model training, and inference services. Concept validation is performed using the Alibaba Cluster Trace Dataset and Evidently AI Drift Detection Dataset, demonstrating the practicality and applicability of the proposed telemetry primitives in real-world scenarios. The study further shows that lineage-aware correlation can reveal operational dependencies and failure propagation patterns that remain undetected by conventional component-level monitoring approaches. In addition, the framework defines a tool-agnostic event format and supports machine learning–based incident classification for common cross-stage failure modes, including training-serving skew and prediction degradation caused by data freshness issues. An incremental adoption strategy is proposed to facilitate implementation, beginning with high-impact production models and expanding according to demonstrated operational value.
Keywords
References
S. Shankar and A. Parameswaran, "Towards observability for production machine learning pipelines," arXiv Prepr. arXiv2108.13557, 2021.
M. Schmuck, "Cultivating Data Observability As The Next Frontier Of Data Engineering," J. Public Adm. Financ. Law, no. 30 Special, pp. 212–224, 2023.
D. Sculley et al., "Hidden technical debt in machine learning systems," Adv. Neural Inf. Process. Syst., vol. 28, 2015.
N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, "Data management challenges in production machine learning," in Proceedings of the 2017 ACM international conference on management of data, 2017, pp. 1723–1726.
S. Amershi et al., "Software engineering for machine learning: A case study," in 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2019, pp. 291–300.
M. Zaharia et al., "Accelerating the machine learning lifecycle with MLflow," IEEE Data Eng—bull., vol. 41, no. 4, pp. 39–45, 2018.
L. P. Rongali, "Performance Overhead and Optimization Strategies in Opentelemetry," Authorea Prepr., 2025.
P. L. Foalem, F. Khomh, and H. Li, "Studying logging practice in machine learning-based applications," Inf. Softw. Technol., vol. 170, p. 107450, 2024.
C. Olston et al., "Tensorflow-serving: Flexible, high-performance ML serving," arXiv Prepr. arXiv1712.06139, 2017.
D. Crankshaw et al., "The missing piece in complex analytics: Low latency, scalable model management and serving with velox," arXiv Prepr. arXiv1409.3809, 2014.
D. Baylor et al., "Tfx: A TensorFlow-based production-scale machine learning platform," in Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, 2017, pp. 1387–1395.
B. H. Sigelman et al., "Dapper, a large-scale distributed systems tracing infrastructure," 2010.
J. Kaldor et al., "Canopy: An end-to-end performance tracing and analysis system," in Proceedings of the 26th symposium on operating systems principles, 2017, pp. 34–50.
J. Shen et al., "Network-centric distributed tracing with deepflow: Troubleshooting your microservices in zero code," in Proceedings of the ACM SIGCOMM 2023 Conference, 2023, pp. 420–437.
J. Mace, R. Roelke, and R. Fonseca, "Pivot tracing: Dynamic causal monitoring for distributed systems," ACM Trans. Comput. Syst., vol. 35, no. 4, pp. 1–28, 2018.
N. Bangad et al., "A Theoretical Framework for AI-driven data quality monitoring in high-volume data environments," arXiv Prepr. arXiv2410.08576, 2024.
M. Interlandi et al., "Adding data provenance support to Apache Spark," VLDB J., vol. 27, no. 5, pp. 595–615, 2018.
E. Breck, S. Cai, E. Nielsen, M. Salib, and D. Sculley, "The ML test score: A rubric for ML production readiness and technical debt reduction," in 2017 IEEE International Conference on Big Data (Big Data), 2017, pp. 1123–1132.
A. Cloud, "Alibaba Cluster Trace Program," GitHub Repository. [Online]. Available: https://github.com/alibaba/clusterdata
R. R. Sambasivan, I. Shafer, J. Mace, B. H. Sigelman, R. Fonseca, and G. R. Ganger, "Principled workflow-centric tracing of distributed systems," in Proceedings of the Seventh ACM Symposium on Cloud Computing, 2016, pp. 401–414.
E. AI, "Evidently AI Open Source ML Monitoring Examples and Datasets," GitHub Repository. [Online]. Available: https://github.com/evidentlyai/evidently/tree/main/examples
N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich, "Data lifecycle challenges in production machine learning: a survey," ACM Sigmod Rec., vol. 47, no. 2, pp. 17–28, 2018.
DOI: https://doi.org/10.52088/ijesty.v6i1.1824
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Shankar das Boddu




























