Preview

Intelligent transport

Advanced search

Comparative Analysis of Methods for Improving RAG System Accuracy on Specialized Corpora of Regulatory and Technical Documentation

https://doi.org/10.24412/3033-6007-2026-339-4-22

Abstract

The digital transformation of industry drives a growing demand for intelligent decision support systems capable of efficiently extracting knowledge from regulatory and technical documentation. This paper examines Retrieval-Augmented Generation (RAG) technology applied to a specialized corpus of railway regulatory documents. A comparative analysis is conducted for several methods aimed at improving RAG accuracy, including query processing strategies (HyDE, Step-Back, Multi-Query, Decomposition, Pseudo-Relevance Feedback, Recursive Refinement), two-stage retrieval with a cross-encoder, and fine-tuning of embedding models on synthetic question-regulatory clause pairs. The experiments demonstrate that domain-specific fine-tuning of embedding models yields the most significant improvement in retrieval quality. Cross-encoder reranking also provides a positive effect, particularly for models with initially less accurate ranking. At the same time, advanced query processing strategies do not lead to a substantial improvement in retrieval quality compared to the baseline approach. The constructed document corpus, question sets, and source code are made publicly available.

About the Authors

A. A. Agafonov
St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS)
Russian Federation

Postgraduate Student, Junior Researcher



A. V. Ponomarev
St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS)
Russian Federation

Candidate of Technical Sciences, Associate Professor, Senior Researcher



References

1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrievalaugmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

2. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., et al. (2023). Retrieval-augmented generation for large language models: A survey (arXiv:2312.10997). arXiv. https://doi.org/10.48550/a rXiv.2312.10997

3. Vrdoljak, J., Boban, Z., Vilović, M., Kumrić, M., & Božić, J. (2025). A review of large language models in medical education, clinical decision support, and healthcare administration. Healthcare, 13(6), Article 603. https://doi.org/10.3390/healthcare13060603

4. Vizniuk, A., Diachenko, G., Laktionov, I., Siwocha, A., Xiao, M., & Smoląg, J. (2025). A comprehensive survey of retrieval-augmented large language models for decision making in agriculture: Unsolved problems and research opportunities. Journal of Artificial Intelligence and Soft Computing Research, 15(2), 115–146. https://doi.org/10.2478/jaiscr-2025-0007

5. Ershov, M. A. (2025). Intelligent DSS for competitiveness assessment in e-commerce: Methodology of RAG-system integration. Ekonomika stroitelstva, (11), 518–521. (in Russian)

6. Samofalov, G. S. (2026). Integration of large language models (LLM) and RAG into decision support systems (DSS) for pricing and demand management in retail. Internauka, (16-1), 38–42. (in Russian)

7. Cioplea, R. (2026). Memory in the age of AI: How LLMs remember, forget, and leak (Technical Report ver. 3.0). https://doi.org/10.17605/OSF.IO/9WXG3

8. Es, S., James, J., Anke, L. E., & Schockaert, S. (2024). RAGAS: Automated evaluation of retrieval-augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). https://doi.org/10.18653/v1/2024.eacl-demo.16

9. Gao, L., Ma, X., Lin, J., & Callan, J. (2023). Precise zero-shot dense retrieval without relevance labels. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1762–1777). https://doi.org/10.18653/v1/2023.acl-long.99

10. Zheng, H. S., Mishra, S., Chen, X., Cheng, H.-T., Chi, E. H., Le, Q. V., & Zhou, D. (2024). Take a step back: Evoking reasoning via abstraction in large language models. In International Conference on Learning Representations (ICLR 2024).

11. Li, Z., et al. (2024). DMQR-RAG: Diverse multi-query rewriting for RAG (arXiv:2411.13154). arXiv. https://doi.org/10.48550/arXiv.2411.13154

12. Ammann, P. J. L., Golde, J., & Akbik, A. (2025). Question decomposition for retrievalaugmented generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop) (pp. 497–507). https: //doi.org/10.18653/v1/2025.acl-srw.32

13. Salton, G., & Buckley, C. (1990). Improving retrieval performance by relevance feedback. Journal of the American Society for Information Science, 41(4), 288–297.

14. Mackie, I., Chatterjee, S., & Dalton, J. (2023). Generative relevance feedback with large language models. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2026–2031).

15. Wang, Y. (2024). MaFeRw: Query rewriting with multi-aspect feedbacks for retrieval-augmented large language models (arXiv:2408.17072). arXiv. https://doi.org/10.48550/arXiv.2408.17072

16. Nogueira, R., & Cho, K. (2019). Passage re-ranking with BERT (arXiv:1901.04085). arXiv. https://doi.org/10.48550/arXiv.1901.04085

17. Chen, J., et al. (2024). M3-Embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation (arXiv:2402.03216). arXiv. https://doi.org/10.48550/arXiv.2402.03216

18. Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982–3992). https://doi.org/10.18653/v1/D19-1410

19. Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., et al. (2022). Text embeddings by weakly-supervised contrastive pre-training (arXiv:2212.03533). arXiv. https: //doi.org/10.48550/arXiv.2212.03533

20. Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2022). MTEB: Massive text embedding benchmark (arXiv:2210.07316). arXiv. https://doi.org/10.48550/arXiv.2210.07316

21. Snegirev, A., et al. (2024). The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design (arXiv:2408.12503). arXiv. https://doi.org/10.48550/arXiv.2408.12503

22. Wang, L., et al. (2024). Multilingual E5 text embeddings: A technical report (arXiv:2402.05672). arXiv. https://doi.org/10.48550/arXiv.2402.05672

23. Parshin, K. A., & Podgornyy, M. S. (2019). A study of the application of railway industry term features in the formation of a classifier. Teoriya i praktika sovremennoy nauki, (3), 229–232. (in Russian)

24. Yang, A., et al. (2025). Qwen3 technical report (arXiv:2505.09388). arXiv. https://doi.org/10.48550/arXiv.2505.09388


Review

For citations:


Agafonov A.A., Ponomarev A.V. Comparative Analysis of Methods for Improving RAG System Accuracy on Specialized Corpora of Regulatory and Technical Documentation. Intelligent transport. 2026;10(3(39)):4-22. https://doi.org/10.24412/3033-6007-2026-339-4-22

Views: 98

JATS XML


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 3033-6007 (Online)