Article Open Access

When Pipelines Break Silently: The Case for Self-Adaptive Data Cleaning

Siddharth Kumar Choudhary

Abstract


Schema changes have transitioned from occasional maintenance events to a routine operational condition in modern data engineering environments. Yet, the theoretical foundations for handling it at the operator level remain underdeveloped. Upstream teams modify field names, remove columns, alter data types, and restructure nested objects as a natural consequence of product iteration and infrastructure migration. At the same time, downstream cleaning logic continues executing against structural assumptions that no longer hold. The result is pipelines that survive schema change while silently degrading the quality of their output in ways that standard monitoring frameworks do not immediately surface. Recent advances in schema evolution cataloging, data quality optimization, provenance tracking, and CI/CD governance have each strengthened individual components of the pipeline reliability problem. Still, they have treated these concerns as neighboring domains rather than as parts of a unified theory. The concept of operator robustness to schema evolution addresses this fragmentation directly by defining robustness not as functional continuity alone, but as the joint preservation of function and data quality after structural change. Mapping the cleaning operator's sensitivity to elementary schema modification operations yields a decision framework that transforms reactive maintenance into an analyzable adaptation problem. Embedding that framework within a self-adaptive architectural loop—monitoring, analysis, planning, and execution over shared knowledge—establishes the conceptual infrastructure for pipelines that respond to structural uncertainty autonomously rather than depending on manual repair. The forward agenda includes tri-objective optimization across quality, latency, and compute cost, provenance-guided operator repair, contract-aware adaptation policy, and empirical benchmarking against realistic long-running pipelines with controlled schema injection.


Keywords


Schema Evolution, Operator Robustness, Self-Adaptive Pipelines, Data Quality Optimization, Provenance-Guided Repair

References


S. K. Singu, "Designing scalable data engineering pipelines using Azure and Databricks," ESP J. Eng. & Technol. Adv., vol. 1, no. 2, pp. 176–187, 2021.

D. Sebastian, "Modernizing Credit Risk with Data Mesh: A Large Bank's Transformation to Real-Time Credit Intelligence," J. Comput. Sci. Technol. Stud., vol. 7, no. 10, pp. 650–664, 2025.

B. O. Mkpa, "The Role of Pharmaceutical Quality Control in Preventing Public Health Risks in the United States," Am. J. Adv. Technol. Eng. Solut., vol. 1, no. 02, pp. 135–172, 2025.

G. Papadimitriou, D. Gizopoulos, H. D. Dixit, and S. Sankar, "Silent data corruptions: The stealthy saboteurs of digital integrity," in 2023 IEEE 29th International Symposium on On-Line Testing and Robust System Design (IOLTS), 2023, pp. 1–7.

A. Asudeh, "Student Travel Support for the 51st International Conference on Very Large Databases 2025," NSF Award Number 2525516. Dir. Comput. Inf. Sci. Eng., vol. 25, no. 2525516, p. 25516, 2025.

J. Bogner, R. Verdecchia, and I. Gerostathopoulos, "Characterizing technical debt and antipatterns in AI-based systems: A systematic mapping study," in 2021 IEEE/ACM International Conference on Technical Debt (TechDebt), 2021, pp. 64–73.

E. Gkintoni, H. Antonopoulou, A. Sortwell, and C. Halkiopoulos, "Challenging cognitive load theory: The role of educational neuroscience and artificial intelligence in redefining learning efficacy," Brain Sci., vol. 15, no. 2, p. 203, 2025.

Z. Brahmia, F. Grandi, and B. Oliboni, "A literature review on schema evolution in databases," Comput. Open, vol. 2, p. 2430001, 2024.

Z. Brahmia, F. Grandi, and B. Oliboni, "Schema Versioning in Databases: A Literature Review," Comput. Open, vol. 2, p. 2430002, 2024.

V. Restat, M. Klettke, and U. Störl, "Towards an end-to-end data quality optimizer," in 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW), 2024, pp. 262–266.

V. Restat, I. Diestelkämper, M. Klettke, and U. Störl, "FONDUE—Fine-Tuned Optimization: Nurturing Data Usability & Efficiency," J. Big Data, vol. 12, no. 1, p. 131, 2025.

S. Strasser and M. Klettke, "Transparent data preprocessing for machine learning," in Proceedings of the 2024 Workshop on Human-In-the-Loop Data Analytics, 2024, pp. 1–6.

V. Restat and U. Störl, "ALPINE: Abstract Language for Pipeline Integration and Execution," in Datenbanksysteme für Business, Technologie und Web-Workshopband (BTW 2025), 2025, pp. 207–217.

A. Chapman, L. Lauro, P. Missier, and R. Torlone, "Supporting better insights of data science pipelines with fine-grained provenance," ACM Trans. Database Syst., vol. 49, no. 2, pp. 1–42, 2024.

S. Schelter, S. Guha, and S. Grafberger, "Automated provenance-based screening of ML data preparation pipelines," Datenbank-Spektrum, vol. 24, no. 3, pp. 187–196, 2024.

F. Bachinger, L. Ehrlinger, G. Kronberger, and W. Wöß, "Data validation utilizing expert knowledge and shape constraints," ACM J. Data Inf. Qual., vol. 16, no. 2, pp. 1–27, 2024.

H. Yang, Z. Xu, S. Yudin, and A. Davidson, "Unlocking the Power of CI/CD for Data Pipelines in Distributed Data Warehouses," Proc. VLDB Endow., vol. 18, no. 12, pp. 4887–4895, 2025.

K. Shivashankar, G. Al Hajj, and A. Martini, "Maintainability and scalability in machine learning: challenges and solutions," ACM Comput. Surv., vol. 57, no. 12, pp. 1–36, 2025.

S. D. Fu and X. Chen, "Compound Schema Registry," arXiv Prepr. arXiv2406.11227, 2024.

B. Pernici et al., "Sustainable quality in data preparation," ACM J. Data Inf. Qual., vol. 17, no. 4, pp. 1–33, 2025.




DOI: https://doi.org/10.52088/ijesty.v6i1.1829

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Siddharth Kumar Choudhary

International Journal of Engineering, Science, and Information Technology (IJESTY) eISSN 2775-2674