Comparative Analysis of Methods for Improving RAG System Accuracy on Specialized Corpora of Regulatory and Technical Documentation
https://doi.org/10.24412/3033-6007-2026-339-4-22
Abstract
The digital transformation of industry drives a growing demand for intelligent decision support systems capable of efficiently extracting knowledge from regulatory and technical documentation. This paper examines Retrieval-Augmented Generation (RAG) technology applied to a specialized corpus of railway regulatory documents. A comparative analysis is conducted for several methods aimed at improving RAG accuracy, including query processing strategies (HyDE, Step-Back, Multi-Query, Decomposition, Pseudo-Relevance Feedback, Recursive Refinement), two-stage retrieval with a cross-encoder, and fine-tuning of embedding models on synthetic question-regulatory clause pairs. The experiments demonstrate that domain-specific fine-tuning of embedding models yields the most significant improvement in retrieval quality. Cross-encoder reranking also provides a positive effect, particularly for models with initially less accurate ranking. At the same time, advanced query processing strategies do not lead to a substantial improvement in retrieval quality compared to the baseline approach. The constructed document corpus, question sets, and source code are made publicly available.
About the Authors
A. A. AgafonovRussian Federation
Postgraduate Student, Junior Researcher
A. V. Ponomarev
Russian Federation
Candidate of Technical Sciences, Associate Professor, Senior Researcher
References
1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrievalaugmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.
2. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., et al. (2023). Retrieval-augmented generation for large language models: A survey (arXiv:2312.10997). arXiv. https://doi.org/10.48550/a rXiv.2312.10997
3. Vrdoljak, J., Boban, Z., Vilović, M., Kumrić, M., & Božić, J. (2025). A review of large language models in medical education, clinical decision support, and healthcare administration. Healthcare, 13(6), Article 603. https://doi.org/10.3390/healthcare13060603
4. Vizniuk, A., Diachenko, G., Laktionov, I., Siwocha, A., Xiao, M., & Smoląg, J. (2025). A comprehensive survey of retrieval-augmented large language models for decision making in agriculture: Unsolved problems and research opportunities. Journal of Artificial Intelligence and Soft Computing Research, 15(2), 115–146. https://doi.org/10.2478/jaiscr-2025-0007
5. Ershov, M. A. (2025). Intelligent DSS for competitiveness assessment in e-commerce: Methodology of RAG-system integration. Ekonomika stroitelstva, (11), 518–521. (in Russian)
6. Samofalov, G. S. (2026). Integration of large language models (LLM) and RAG into decision support systems (DSS) for pricing and demand management in retail. Internauka, (16-1), 38–42. (in Russian)
7. Cioplea, R. (2026). Memory in the age of AI: How LLMs remember, forget, and leak (Technical Report ver. 3.0). https://doi.org/10.17605/OSF.IO/9WXG3
8. Es, S., James, J., Anke, L. E., & Schockaert, S. (2024). RAGAS: Automated evaluation of retrieval-augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). https://doi.org/10.18653/v1/2024.eacl-demo.16
9. Gao, L., Ma, X., Lin, J., & Callan, J. (2023). Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1762–1777). https://doi.org/10.18653/v1/2023.acl-long.99
10. Zheng, H. S., Mishra, S., Chen, X., Cheng, H.-T., Chi, E. H., Le, Q. V., & Zhou, D. (2024). Take a step back: Evoking reasoning via abstraction in large language models. In International Conference on Learning Representations (ICLR 2024).
11. Li, Z., et al. (2024). DMQR-RAG: Diverse multi-query rewriting for RAG (arXiv:2411.13154). arXiv. https://doi.org/10.48550/arXiv.2411.13154
12. Ammann, P. J. L., Golde, J., & Akbik, A. (2025). Question decomposition for retrievalaugmented generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop) (pp. 497–507). https: //doi.org/10.18653/v1/2025.acl-srw.32
13. Salton, G., & Buckley, C. (1990). Improving retrieval performance by relevance feedback. Journal of the American Society for Information Science, 41(4), 288–297.
14. Mackie, I., Chatterjee, S., & Dalton, J. (2023). Generative relevance feedback with large language models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2026–2031).
15. Wang, Y. (2024). MaFeRw: Query rewriting with multi-aspect feedbacks for retrieval-augmented large language models (arXiv:2408.17072). arXiv. https://doi.org/10.48550/arXiv.2408.17072
16. Nogueira, R., & Cho, K. (2019). Passage re-ranking with BERT (arXiv:1901.04085). arXiv. https://doi.org/10.48550/arXiv.1901.04085
17. Chen, J., et al. (2024). M3-Embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation (arXiv:2402.03216). arXiv. https://doi.org/10.48550/arXiv.2402.03216
18. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). https://doi.org/10.18653/v1/D19-1410
19. Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., et al. (2022). Text embeddings by weakly-supervised contrastive pre-training (arXiv:2212.03533). arXiv. https: //doi.org/10.48550/arXiv.2212.03533
20. Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2022). MTEB: Massive text embedding benchmark (arXiv:2210.07316). arXiv. https://doi.org/10.48550/arXiv.2210.07316
21. Snegirev, A., et al. (2024). The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design (arXiv:2408.12503). arXiv. https://doi.org/10.48550/arXiv.2408.12503
22. Wang, L., et al. (2024). Multilingual E5 text embeddings: A technical report (arXiv:2402.05672). arXiv. https://doi.org/10.48550/arXiv.2402.05672
23. Parshin, K. A., & Podgornyy, M. S. (2019). A study of the application of railway industry term features in the formation of a classifier. Teoriya i praktika sovremennoy nauki, (3), 229–232. (in Russian)
24. Yang, A., et al. (2025). Qwen3 technical report (arXiv:2505.09388). arXiv. https://doi.org/10.48550/arXiv.2505.09388
Review
For citations:
Agafonov A.A., Ponomarev A.V. Comparative Analysis of Methods for Improving RAG System Accuracy on Specialized Corpora of Regulatory and Technical Documentation. Intelligent transport. 2026;10(3(39)):4-22. https://doi.org/10.24412/3033-6007-2026-339-4-22
JATS XML




