Article Open Access

Architecting Reliable Knowledge Retrieval Systems Using Large Language Models

Bharat Kumar Reddy Karumuri

Abstract


The increasing deployment of large language models (LLMs) in enterprise environments creates significant reliability challenges, including hallucination, factual inconsistency, limited knowledge traceability, uncertainty, and operational inefficiency. This study develops a literature-based architectural framework for reliable knowledge retrieval systems that separates external knowledge management from LLM-based reasoning and generation. The framework synthesizes key architectural mechanisms, including knowledge representation, hybrid retrieval, reranking, evidence selection, context construction, response verification, provenance tracking, uncertainty handling, guardrails, and computational efficiency. The resulting architecture organizes these mechanisms into coordinated layers that regulate the flow of external evidence from knowledge sources to generated responses while supporting traceability, evidence grounding, and controlled abstention when sufficient evidence is unavailable. The architectural synthesis further identifies complementary strategies for enterprise deployment, including semantic coaching, model routing, and human oversight, to balance reliability, scalability, responsiveness, and operational cost. The analysis indicates that reliable LLM deployment should be addressed as an end-to-end architectural challenge rather than solely as a model-performance problem. In this perspective, knowledge access, evidence quality, retrieval accuracy, verification, provenance, uncertainty management, and governance must function as integrated system components. The proposed framework provides a structured foundation for designing maintainable, auditable, reliable knowledge retrieval systems capable of supporting enterprise applications and other high-stakes environments. By explicitly separating knowledge acquisition, retrieval, reasoning, verification, and governance functions, the framework also facilitates modular implementation, systematic evaluation, and continuous improvement. Consequently, it offers practical architectural guidance for organizations seeking to deploy LLM-based systems while maintaining factual reliability, operational control, transparency, and accountability across evolving enterprise knowledge environments

Keywords


Knowledge Representation, Retrieval-Augmented Generation, Uncertainty Quantification, Enterprise Governance, Human-AI Collaboration

References


A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, ?. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems, vol. 30, 2017. doi: https://doi.org/10.48550/arXiv.1706.03762.

R. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021. doi: https://doi.org/10.48550/arXiv.2108.07258.

Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, Art. no. 248, 2023. doi: https://doi.org/10.1145/3571730.

P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020. doi: https://doi.org/10.48550/arXiv.2005.11401.

V. Karpukhin et al., “Dense passage retrieval for open-domain question answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 6769–6781, 2020. doi: https://doi.org/10.18653/v1/2020.emnlp-main.550.

X. Ma et al., “Pre-train a discriminative text encoder for dense retrieval via contrastive span prediction,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1360–1372, 2022. doi: https://doi.org/10.1145/3477495.3531772.

H. Trivedi et al., “Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,” arXiv preprint arXiv:2212.10509, 2023. doi: https://doi.org/10.48550/arXiv.2212.10509.

A. Asai et al., “Self-RAG: Learning to retrieve, generate, and critique through self-reflection,” arXiv preprint arXiv:2310.11511, 2023. doi: https://doi.org/10.48550/arXiv.2310.11511.

N. F. Liu et al., “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 157–173, 2024. doi: https://doi.org/10.1162/tacl_a_00638.

Y. Geifman and R. El-Yaniv, “Selective classification for deep neural networks,” Advances in Neural Information Processing Systems, vol. 30, 2017. doi: https://doi.org/10.48550/arXiv.1705.08500.

S. Kadavath et al., “Language models (mostly) know what they know,” arXiv preprint arXiv:2207.05221, 2022. doi: https://doi.org/10.48550/arXiv.2207.05221.

H. Inan et al., “Llama Guard: LLM-based input-output safeguard for human-AI conversations,” arXiv preprint arXiv:2312.06674, 2023. doi: https://doi.org/10.48550/arXiv.2312.06674.

Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” arXiv preprint arXiv:2211.17192, 2022. doi: https://doi.org/10.48550/arXiv.2211.17192.

O. Honovich et al., “TRUE: Re-evaluating factual consistency evaluation,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3905–3920, 2022. doi: https://doi.org/10.18653/v1/2022.naacl-main.287.

J. D. Lee and K. A. See, “Trust in automation: Designing for appropriate reliance,” Human Factors, vol. 46, no. 1, pp. 50–80, 2004. doi: https://doi.org/10.1518/hfes.46.1.50_30392.

W. Sun et al., “Is ChatGPT good at search? Investigating large language models as re-ranking agents,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 14918–14937, 2023. doi: https://doi.org/10.18653/v1/2023.emnlp-main.923.

S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. C. Park, “Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 7036–7050, 2024. doi: https://doi.org/10.18653/v1/2024.naacl-long.389.

S. Es, J. James, L. Espinosa-Anke, and S. Schockaert, “RAGAs: Automated evaluation of retrieval augmented generation,” in Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pp. 150–158, 2024. doi: https://doi.org/10.18653/v1/2024.eacl-demo.16.

D. Ru et al., “RAGChecker: A fine-grained framework for diagnosing retrieval-augmented generation,” Advances in Neural Information Processing Systems, vol. 37, 2024. doi: https://doi.org/10.48550/arXiv.2408.08067.

S. Mao et al., “RaFe: Ranking feedback improves query rewriting for RAG,” Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 884–901, 2024. doi: https://doi.org/10.18653/v1/2024.findings-emnlp.49.

Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, Art. no. 248, 2023. doi: https://doi.org/10.1145/3571730.

L. S. Hartono, E. I. Setiawan, and V. Singh, “Retrieval augmented generation-based chatbot for prospective and current university students,” International Journal of Engineering, Science and Information Technology, vol. 5, no. 3, pp. 268–277, 2025. doi: https://doi.org/10.52088/ijesty.v5i3.951.

Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, M. Wang, and H. Wang, “Retrieval-augmented generation for large language models: A survey,” arXiv preprint arXiv:2312.10997, 2023. doi: https://doi.org/10.48550/arXiv.2312.10997.

N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992, 2019. doi: https://doi.org/10.18653/v1/D19-1410.

S. Robertson and H. Zaragoza, “The probabilistic relevance framework: BM25 and beyond,” Foundations and Trends in Information Retrieval, vol. 3, no. 4, pp. 333–389, 2009. doi: https://doi.org/10.1561/1500000019.

R. Nogueira and K. Cho, “Passage re-ranking with BERT,” arXiv preprint arXiv:1901.04085, 2019. doi: https://doi.org/10.48550/arXiv.1901.04085.

O. Khattab and M. Zaharia, “ColBERT: Efficient and effective passage search via contextualized late interaction over BERT,” in Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 39–48, 2020. doi: https://doi.org/10.1145/3397271.3401075.

Z. Jiang, F. F. Xu, L. Gao, Z. Sun, Q. Liu, J. Dwivedi-Yu, Y. Yang, J. Callan, and G. Neubig, “Active retrieval augmented generation,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7969–7992, 2023. doi: https://doi.org/10.18653/v1/2023.emnlp-main.495.

S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “ReAct: Synergizing reasoning and acting in language models,” in International Conference on Learning Representations (ICLR), 2023. doi: https://doi.org/10.48550/arXiv.2210.03629.

S. Lin, J. Hilton, and O. Evans, “Teaching models to express their uncertainty in words,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 1136–1154, 2022. doi: https://doi.org/10.1162/tacl_a_00431.

K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, “Retrieval augmentation reduces hallucination in conversation,” in Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3784–3803, 2021. doi: https://doi.org/10.18653/v1/2021.findings-emnlp.320.

R. Chandra, R. Bansal, and K. Lulla, “Benchmarking techniques for real-time evaluation of LLMs in production systems,” International Journal of Engineering, Science and Information Technology, vol. 5, no. 3, pp. 363–372, 2025. doi: https://doi.org/10.52088/ijesty.v5i3.955.

H. Yu, A. Gan, K. Zhang, S. Tong, Q. Liu, and Z. Liu, “RAGChecker: A fine-grained framework for diagnosing retrieval-augmented generation,” Advances in Neural Information Processing Systems, vol. 37, 2024. doi: https://doi.org/10.48550/arXiv.2408.08067.

K. Peffers, T. Tuunanen, M. A. Rothenberger, and S. Chatterjee, “A design science research methodology for information systems research,” Journal of Management Information Systems, vol. 24, no. 3, pp. 45–77, 2007. doi: https://doi.org/10.2753/MIS0742-1222240302.

B. Bohnet, V. Q. Tran, P. Verga, et al., “Attributed question answering: Evaluation and modeling for attributed large language models,” arXiv preprint arXiv:2212.08037, 2022. doi: https://doi.org/10.48550/arXiv.2212.08037.

X. Ma, Y. Gong, P. He, H. Zhao, and N. Duan, “Query rewriting for retrieval-augmented large language models,” arXiv preprint arXiv:2305.14283, 2023. doi: https://doi.org/10.48550/arXiv.2305.14283.

G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 874–880, 2021. doi: https://doi.org/10.18653/v1/2021.eacl-main.74.

O. Ram, Y. Levine, I. Dalmedigos, D. Muhlgay, A. Shashua, K. Leyton-Brown, and Y. Shoham, “In-context retrieval-augmented language models,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 1316–1331, 2023. doi: https://doi.org/10.1162/tacl_a_00605.

S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi, “FActScore: Fine-grained atomic evaluation of factual precision in long form text generation,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12076–12100, 2023. doi: https://doi.org/10.18653/v1/2023.emnlp-main.566.

T. Gao, H. Yen, J. Yu, and D. Chen, “Enabling large language models to generate text with citations,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023. doi: https://doi.org/10.18653/v1/2023.emnlp-main.398.

H. Rashkin, V. Nikolaev, M. Lamm, L. Aroyo, M. Collins, D. Das, S. Petrov, and D. Reitter, “Measuring attribution in natural language generation models,” Computational Linguistics, vol. 49, no. 4, pp. 777–840, 2023. doi: https://doi.org/10.1162/coli_a_00486




DOI: https://doi.org/10.52088/ijesty.v6i3.1908

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Bharat Kumar Reddy Karumuri

International Journal of Engineering, Science, and Information Technology (IJESTY) eISSN 2775-2674