Provisioning AI-Ready Infrastructure at Scale: Engineering Considerations for Infrastructure Architects
Abstract
Keywords
References
L. Alzubaidi, J. Zhang, A. J. Humaidi, A. Al-Dujaili, Y. Duan, O. Al-Shamma, et al., “Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions,” Journal of Big Data, vol. 8, no. 1, art. 53, 2021. https://doi.org/10.1186/s40537-021-00444-8
J. Shalf, “The future of computing beyond Moore’s Law,” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 2020. https://doi.org/10.1098/rsta.2019.0061
Z. Chen, W. Quan, M. Wen, J. Fang, J. Yu, et al., “Deep learning research and development platform: Characterizing and scheduling with QoS guarantees on GPU clusters,” IEEE Transactions on Parallel and Distributed Systems, 2019. https://ieeexplore.ieee.org/abstract/document/8778770/
R. Kaur, A. Asad, S. Al Abdul Wahid, and F. Mohammadi, “A survey of advancements in scheduling techniques for efficient deep learning computations on GPUs,” Electronics, vol. 14, no. 5, art. 1048, 2025. https://doi.org/10.3390/electronics14051048
Z. Ye, W. Gao, Q. Hu, P. Sun, X. Wang, Y. Luo, T. Zhang, and Y. Wen, “Deep learning workload scheduling in GPU datacenters: A survey,” ACM Computing Surveys, vol. 56, no. 6, art. 146, pp. 1–38, 2024. https://doi.org/10.1145/3638757
T. Fu, Z. Yang, Z. Ye, C. Ma, Y. Han, Y. Luo, X. Wang, Z. Wang, and X. Lei, “A survey on the scheduling of DL and LLM training jobs in GPU clusters,” Chinese Journal of Electronics, vol. 34, no. 3, pp. 881–905, 2025. https://doi.org/10.23919/cje.2024.00.070
Q. Weng, W. Xiao, Y. Yu, W. Wang, C. Wang, J. He, et al., “MLaaS in the wild: Workload analysis and scheduling in large-scale heterogeneous GPU clusters,” in Proc. 19th USENIX Symp. Networked Systems Design and Implementation (NSDI), 2022. https://www.usenix.org/conference/nsdi22/presentation/weng
Q. Hu, P. Sun, S. Yan, Y. Wen, and T. Zhang, “Characterization and prediction of deep learning workloads in large-scale GPU datacenters,” in Proc. Int. Conf. High Performance Computing, Networking, Storage and Analysis (SC ’21), 2021, art. 104, pp. 1–15. https://doi.org/10.1145/3458817.3476223
M. Wang, C. Meng, G. Long, C. Wu, et al., “Characterizing deep learning training workloads on Alibaba-PAI,” in Proc. IEEE Int. Symp. Workload Characterization (IISWC), 2019. https://ieeexplore.ieee.org/abstract/document/9042047/
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proc. 29th ACM Symp. Operating Systems Principles (SOSP), 2023. https://doi.org/10.1145/3600006.3613165
H. Hu, Y. Wen, T.-S. Chua, and X. Li, “Toward scalable systems for big data analytics: A technology tutorial,” IEEE Access, 2014. https://doi.org/10.1109/ACCESS.2014.2332453
A. A. Awan, A. Jain, C.-H. Chu, H. Subramoni, et al., “Communication profiling and characterization of deep-learning workloads on clusters with high-performance interconnects,” IEEE Micro, 2019. https://ieeexplore.ieee.org/abstract/document/8887216/
E. García-Martín, C. F. Rodrigues, G. Riley, and H. Grahn, “Estimation of energy consumption in machine learning,” Journal of Parallel and Distributed Computing, 2019. https://doi.org/10.1016/j.jpdc.2019.07.007
B. Varghese and R. Buyya, “Next generation cloud computing: New trends and research directions,” Future Generation Computer Systems, 2018. https://doi.org/10.1016/j.future.2017.09.020
R. Boutaba, M. A. Salahuddin, N. Limam, S. Ayoubi, N. Shahriar, F. Estrada-Solano, and O. M. Caicedo, “A comprehensive survey on machine learning for networking: Evolution, applications and research opportunities,” Journal of Internet Services and Applications, vol. 9, no. 1, art. 16, 2018. https://doi.org/10.1186/s13174-018-0087-2
D. Kreuzberger, N. Kühl, and S. Hirschl, “Machine learning operations (MLOps): Overview, definition, and architecture,” IEEE Access, vol. 11, pp. 31866–31879, 2023. https://doi.org/10.1109/ACCESS.2023.3262138
G. Nguyen, S. Dlugolinský, M. Bobák, V. Tran, Á. López García, I. Heredia, P. Malík, and L. Hluchý, “Machine learning and deep learning frameworks and libraries for large-scale data mining: A survey,” Artificial Intelligence Review, 2019. https://doi.org/10.1007/s10462-018-09679-z
M. Jeon, S. Venkataraman, A. Phanishayee, J. Qian, W. Lee, and S. Muddu, “Analysis of large-scale multi-tenant GPU clusters for DNN training workloads,” in Proc. USENIX Annual Technical Conference (USENIX ATC), 2019. https://www.usenix.org/conference/atc19/presentation/jeon
Y. Wang, J. Yu, and Z. Yu, “Resource scheduling techniques in cloud from a view of coordination: A holistic survey,” Frontiers of Information Technology & Electronic Engineering, 2023. https://doi.org/10.1631/FITEE.2100298
M. Aledhari, R. Razzak, R. M. Parizi, and F. Saeed, “Federated learning: A survey on enabling technologies, protocols, and applications,” IEEE Access, 2020. https://doi.org/10.1109/ACCESS.2020.3013541
P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Bennis, A. N. Bhagoji, et al., “Advances and open problems in federated learning,” Foundations and Trends in Machine Learning, vol. 14, nos. 1–2, pp. 1–210, 2021. https://doi.org/10.1561/2200000083
W. Xiao, S. Ren, Y. Li, Y. Zhang, P. Hou, Z. Li, et al., “AntMan: Dynamic scaling on GPU clusters for deep learning,” in Proc. 14th USENIX Symp. Operating Systems Design and Implementation (OSDI), 2020. https://www.usenix.org/conference/osdi20/presentation/xiao
Y. Babuji, A. Woodard, Z. Li, D. S. Katz, B. Clifford, R. Kumar, et al., “Parsl: Pervasive parallel programming in Python,” in Proc. 28th ACM Int. Symp. High-Performance Parallel and Distributed Computing (HPDC), 2019. https://doi.org/10.1145/3307681.3325400
S. S. Gill, H. Wu, P. Patros, C. Ottaviani, P. Arora, et al., “Modern computing: Vision and challenges,” Telematics and Informatics Reports, 2024. https://doi.org/10.1016/j.teler.2024.100116
K. Adimora, S. R. Gundla, and H. Sun, “Machine learning approaches for optimizing high-performance computing scheduling: A comprehensive survey and analysis,” Cluster Computing, 2025. https://doi.org/10.1007/s10586-025-05521-8
Y. K. Dwivedi, L. Hughes, E. Ismagilova, G. Aarts, C. Coombs, T. Crick, et al., “Artificial intelligence (AI): Multidisciplinary perspectives on emerging challenges, opportunities, and agenda for research, practice and policy,” International Journal of Information Management, 2021. https://doi.org/10.1016/j.ijinfomgt.2019.08.002
R. Chab, F. Li, and S. Setia, “Algorithmic techniques for GPU scheduling: A comprehensive survey,” Algorithms, vol. 18, no. 7, art. 385, 2025. https://doi.org/10.3390/a18070385
K. Lulla, R. Chandra, and K. Ranjan, “Factory-Grade Diagnostic Automation for GeForce and Data Centre GPUs,” International Journal of Engineering, Science and Information Technology, vol. 5, no. 3, pp. 537–544, 2025. https://doi.org/10.52088/ijesty.v5i3.1089
K. Senjab, A. Abbas, and N. Ahmed, “A survey of Kubernetes scheduling algorithms,” Journal of Cloud Computing, vol. 12, art. 87, 2023. https://doi.org/10.1186/s13677-023-00471-1
M. M. Najafabadi, F. Villanustre, T. M. Khoshgoftaar, N. Seliya, R. Wald, and E. Muharemagic, “Deep learning applications and challenges in big data analytics,” Journal of Big Data, vol. 2, no. 1, art. 1, 2015. https://doi.org/10.1186/s40537-014-0007-7
J. Goecks, A. Nekrutenko, J. Taylor, and The Galaxy Team, “Galaxy: A comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences,” Genome Biology, vol. 11, no. 8, art. R86, 2010. https://doi.org/10.1186/gb-2010-11-8-r86
M. L. Gambo and A. Almulhem, “Zero trust architecture: A systematic literature review,” Journal of Network and Systems Management, vol. 34, no. 1, art. 25, 2026. https://doi.org/10.1007/s10922-025-09998-x
M. A. Akbar, K. Smolander, S. Mahmood, and A. Alsanad, “Toward successful DevSecOps in software development organizations: A decision-making framework,” Information and Software Technology, vol. 147, art. 106894, 2022. https://doi.org/10.1016/j.infsof.2022.106894
Q. Zhang, L. Cheng, and R. Boutaba, “Cloud computing: State-of-the-art and research challenges,” Journal of Internet Services and Applications, vol. 1, no. 1, pp. 7–18, 2010. https://doi.org/10.1007/s13174-010-0007-6
J. K. Shah and P. Matam, “Towards Self-Healing Cloud Infrastructures: Predictive Maintenance with Reinforcement Learning and Generative Models,” International Journal of Engineering, Science and Information Technology, vol. 5, no. 3, pp. 619–627, 2025. https://doi.org/10.52088/ijesty.v5i3.1185
B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes,” Communications of the ACM, vol. 59, no. 5, pp. 50–57, 2016.
K. Senjab, A. Abbas, and N. Ahmed, “A survey of Kubernetes scheduling algorithms,” Journal of Cloud Computing, vol. 12, art. 87, 2023. https://doi.org/10.1186/s13677-023-00471-1
M. M. Najafabadi, F. Villanustre, T. M. Khoshgoftaar, N. Seliya, R. Wald, and E. Muharemagic, “Deep learning applications and challenges in big data analytics,” Journal of Big Data, vol. 2, no. 1, art. 1, 2015. https://doi.org/10.1186/s40537-014-0007-7
J. Goecks, A. Nekrutenko, J. Taylor, and The Galaxy Team, “Galaxy: A comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences,” Genome Biology, vol. 11, no. 8, art. R86, 2010. https://doi.org/10.1186/gb-2010-11-8-r86
M. L. Gambo and A. Almulhem, “Zero trust architecture: A systematic literature review,” Journal of Network and Systems Management, vol. 34, no. 1, art. 25, 2026. https://doi.org/10.1007/s10922-025-09998-x
M. A. Akbar, K. Smolander, S. Mahmood, and A. Alsanad, “Toward successful DevSecOps in software development organizations: A decision-making framework,” Information and Software Technology, vol. 147, art. 106894, 2022. https://doi.org/10.1016/j.infsof.2022.106894
Q. Zhang, L. Cheng, and R. Boutaba, “Cloud computing: State-of-the-art and research challenges,” Journal of Internet Services and Applications, vol. 1, no. 1, pp. 7–18, 2010. https://doi.org/10.1007/s13174-010-0007-6
Berberi, L., et al. (2025). Machine learning operations landscape: Platforms and tools. Artificial Intelligence Review, 58, Article 167. https://doi.org/10.1007/s10462-025-11164-3
Zhou, N., Zhou, H., & Hoppe, D. (2023). Containerization for high-performance computing systems: Survey and prospects. IEEE Transactions on Software Engineering, 49(4), 2722–2740. https://ieeexplore.ieee.org/document/9985426
Yu, E., et al. (2023). Communication optimization algorithms for distributed deep learning systems: A survey. IEEE Transactions on Parallel and Distributed Systems, 34(12), 3211–3228. https://ieeexplore.ieee.org/document/10275049
Chen, Y., et al. (2024). PeakFS: An ultra-high performance parallel file system via computing-network-storage co-optimization for HPC applications. IEEE Transactions on Parallel and Distributed Systems, 35(12). https://ieeexplore.ieee.org/document/10735121
Schlegel, M., & Sattler, K.-U. (2025). Capturing end-to-end provenance for machine learning pipelines. Information Systems, 122, Article 102347. https://www.sciencedirect.com/science/article/pii/S0306437924001534
Zhang, X., et al. (2025). Communication optimization for distributed training: Architecture, advances, and opportunities. IEEE Network, 39(1), 64–71. https://doi.org/10.1109/MNET.2024.3449276
National Institute of Standards and Technology. (2020). Zero trust architecture (NIST Special Publication 800-207). U.S. Department of Commerce. https://doi.org/10.6028/NIST.SP.800-207
Zanasi, C., et al. (2024). Flexible zero trust architecture for the cybersecurity of industrial IoT infrastructures. Ad Hoc Networks, 157, 103414. https://doi.org/10.1016/j.adhoc.2024.103414
Zeini, A., Lennon, R. G., & Lennon, P. (2023). Securing infrastructure as code (IaC) through DevSecOps: A comprehensive risk management framework. In Proceedings of the 2023 Cyber Research Conference Ireland (Cyber-RCI). https://ieeexplore.ieee.org/document/10671452
P. Nawrocki and M. Smendowski, “A survey of cloud resource consumption optimization methods,” Journal of Grid Computing, vol. 23, art. 5, 2025. https://doi.org/10.1007/s10723-024-09792-0
[53] Setälä, M., & Mikkonen, T. (2025). Containerization in multi-cloud environment: Roles, strategies, challenges, and solutions for effective implementation. ACM Transactions on Software Engineering and Methodology. https://arxiv.org/html/2403.12980v2
Liang, F., et al. (2024). Resource allocation and workload scheduling for large-scale distributed deep learning: A survey. arXiv preprint arXiv:2406.08115. https://arxiv.org/abs/2406.08115
National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.100-
DOI: https://doi.org/10.52088/ijesty.v6i3.1900
Refbacks
- There are currently no refbacks.
Copyright (c) 2026 Hemanth Kumar Gandavarapu






























