Building Compact Scene Graphs Based on a Topological Map for Autonomous Navigation of a Mobile Robot
https://doi.org/10.24412/3033-6007-2026-339-37-52
Abstract
Autonomous navigation of a mobile robot in human-centered environments requires a map that contains not only a geometric model of the environment for path planning, but also information about environmental objects (doors, furniture, office equipment, etc.). Scene graphs provide such a map representation, where nodes correspond to rooms, locations, and objects, and edges encode spatial connectivity or relationships between objects. Most modern scene graph construction methods have high computational complexity, and the graphs they produce are redundant for the purposes of autonomous robot navigation. This paper proposes a method for constructing Compact Scene Graphs (CSG), which is based on the computationally efficient topological mapping method PRISM-TopoMap and the association of semantic objects with locations on the topological map. The resulting scene graph enables route planning to objects via topological map locations and achieves high-precision localization. The proposed method was experimentally evaluated in the Habitat simulation environment. The experimental results demonstrate that the proposed CSG consumes significantly less memory than traditional metric maps and scene graphs, while providing reliable localization on the topological map through association with semantic objects.
About the Authors
K. F. MuravyevRussian Federation
PhD, Research Fellow
V. I. Romanenko
Russian Federation
Student of the Faculty of Computer Science
References
1. Labbé, M., & Michaud, F. (2019). RTAB-Map as an open-source lidar and visual simultaneous localization and mapping library for large-scale and long-term online operation. Journal of Field Robotics, 36(2), 416–446. https://doi.org/10.1002/rob.21831
2. Koide, K., Yokozuka, M., Oishi, S., & Banno, A. (2024). GLIM: 3D range-inertial localization and mapping with GPU-accelerated scan matching factors. Robotics and Autonomous Systems, , Article 104750. https://doi.org/10.1016/j.robot.2024.104750
3. Muravyev, K., & Yakovlev, K. (2022). Evaluation of RGB-D SLAM in large indoor environments. In International Conference on Interactive Collaborative Robotics (pp. 93–104). https://doi.org/10.1007/978-3-031-23609-9_9
4. Muravyev, K., & Yakovlev, K. S. (2023). Evaluation of topological mapping methods in indoor environments. IEEE Access, 11, 132683–132698. https://doi.org/10.1109/ACCESS.2023.
5. Muravyev, K., Melekhin, A., Yudin, D., & Yakovlev, K. (2025). PRISM-TopoMap: Online topological mapping with place recognition and scan matching. IEEE Robotics and Automation Letters, 10(4), 3126–3133. https://doi.org/10.1109/LRA.2025.3541454
6. Kim, N., Kwon, O., Yoo, H., Choi, Y., Park, J., & Oh, S. (2023). Topological semantic graph memory for image-goal navigation. In Proceedings of the 6th Conference on Robot Learning (CoRL) (Vol. 205, pp. 393–402). PMLR.
7. Maggio, D., Chang, Y., Hughes, N., Trang, M., Griffith, D., Dougherty, C., Cristofalo, E., Schmid, L., & Carlone, L. (2024). Clio: Real-time task-driven open-set 3D scene graphs. IEEE Robotics and Automation Letters, 9(10), 8921–8928. https://doi.org/10.1109/LRA.2024.3451395
8. Hughes, N., Chang, Y., & Carlone, L. (2022). Hydra: A real-time spatial perception system for 3D scene graph construction and optimization (arXiv:2201.13360). arXiv. https://doi.org/10.48550/arXiv.2201.13360
9. Peros, S., Delbruel, S., Michiels, S., Joosen, W., & Hughes, D. (2019). Khronos: Middleware for simplified time management in CPS. In Proceedings of the 13th ACM International Conference on Distributed and Event-based Systems (pp. 127–138). https://doi.org/10.1145/3328905.
10. Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Batra, D., & Mottaghi, R. (2019). Habitat: A platform for embodied AI research. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp.–9346). https://doi.org/10.1109/ICCV.2019.00943
11. Allu, S. H., Kadosh, I., Summers, T., & Xiang, Y. (2024). A modular robotic system for autonomous exploration and semantic updating in large-scale indoor environments (arXiv:2409.15493). arXiv. https://doi.org/10.48550/arXiv.2409.15493
12. Carpenter, G. A., & Grossberg, S. (1993). Adaptive resonance theory (Technical Report). Boston University, Center for Adaptive Systems and Department of Cognitive and Neural Systems.
13. Hossain, J., Faridee, A. Z. M., Roy, N., Freeman, J., Gregory, T., & Trout, T. (2024). TopoNav: Topological navigation for efficient exploration in sparse reward environments. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 693–700). https://doi.org/10.1109/IROS58592.2024.10802380
14. Cao, Z., Zhang, Q., Guang, J., Wu, S., Hu, Z., & Liu, J. (2024). SemanticTopoLoop: Semantic loop closure with 3D topological graph based on quadric-level object map. IEEE Robotics and Automation Letters, 9(5), 4257–4264. https://doi.org/10.1109/LRA.2024.3374169
15. Wald, J., Dhamo, H., Navab, N., & Tombari, F. (2020). Learning 3D semantic scene graphs from D indoor reconstructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3961–3970). https://doi.org/10.1109/CVPR42600.2020.00402
16. Zhang, C., Yu, J., Song, Y., & Cai, W. (2021). Exploiting edge-oriented reasoning for 3D pointbased scene graph analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9705–9715). https://doi.org/10.1109/CVPR46437.2021.00958
17. Qi, M., Lv, C., Fu, Z., Zhang, X., & Ma, H. (2026). SGFormer++: Semantic graph transformer for incremental 3D scene graph generation (arXiv:2606.15328). arXiv. https://doi.org/10.48550/arXiv.2606.15328
18. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 8748–8763). PMLR.
19. Koch, S., Vaskevicius, N., Colosi, M., Hermosilla, P., & Ropinski, T. (2024). SGRec3D: Self-supervised 3D scene graph learning via object-level scene reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 3392–3402). https://doi.org/10.1109/WACV57701.2024.00337
20. Tourani, A., Ejaz, S., Bavle, H., Morilla-Cabello, D., Sanchez-Lopez, J. L., & Voos, H. (2025). vS-Graphs: Integrating visual SLAM and situational graphs through multi-level scene understanding (arXiv:2503.01783). arXiv. https://doi.org/10.48550/arXiv.2503.01783
21. Bavle, H., Sánchez-López, J. L., Shaheer, M., Civera, J., & Voos, H. (2023). S-Graphs+: Real-time localization and mapping leveraging hierarchical representations. IEEE Robotics and Automation Letters, 8(8), 4927–4934. https://doi.org/10.1109/LRA.2023.3290512
22. Gu, Q., Kuwajerwala, A., Morin, S., Jatavallabhula, K. M., Sen, B., Agarwal, A., Rivera, C., Paul, W., Ellis, K., Chellappa, R., Gan, C., de Melo, C. M., Tenenbaum, J. B., Torralba, A., Shkurti, F., & Paull, L. (2024). ConceptGraphs: Open-vocabulary 3D scene graphs for perception and planning. In 2024 IEEE International Conference on Robotics and Automation (ICRA) (pp. 5021–5028). https://doi.org/10.1109/ICRA57147.2024.10610243
23. Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Jiang, Q., Li, C., Yang, J., Su, H., Zhu, J., & Zhang, L. (2024). Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. In Computer Vision – ECCV 2024 (Vol. 47, pp. 38–55). Springer. https://doi.org/10.1007/978-3-031-72970-6_3
Review
For citations:
Muravyev K.F., Romanenko V.I. Building Compact Scene Graphs Based on a Topological Map for Autonomous Navigation of a Mobile Robot. Intelligent transport. 2026;10(3(39)):37-52. https://doi.org/10.24412/3033-6007-2026-339-37-52
JATS XML




