融合深度与语义信息的双目视觉SLAM
Binocular Vision SLAM Integrating Depth and Semantic Information
摘要: 针对动态场景下传统视觉SLAM系统因动态物体干扰导致定位精度下降的问题,本文提出一种融合深度特征与语义信息的双目视觉SLAM优化方法。首先,采用轻量化改进的DeepLabV3+语义分割网络,在保证实时性的同时提升分割精度;其次,设计语义与深度一致性融合的动态掩码构建机制:通过语义分割初筛潜在动态区域,结合双目深度图分析时序深度变化,并引入置信度加权评分及时序一致性判断策略,精准区分动态与静态特征点;最后,在ORB-SLAM3框架中集成上述模块,构建包含跟踪、局部建图与回环检测的完整系统。在Euroc和KITTI数据集上的实验表明:与ORB-SLAM3相比,本文方法在室内场景绝对轨迹误差(ATE)降低8%~15%,室外场景降低7%~35%,显著提升了动态环境下的鲁棒性与定位精度。
Abstract: This paper proposes a binocular visual SLAM optimization method that integrates depth features and semantic information, addressing the issue of reduced positioning accuracy in traditional visual SLAM systems due to interference from dynamic objects in dynamic scenes. Firstly, a lightweight improved DeepLabV3 semantic segmentation network is adopted to enhance segmentation accuracy while ensuring real-time performance; secondly, a dynamic mask construction mechanism based on semantic and depth consistency is designed: it initially filters potential dynamic areas through semantic segmentation, analyzes temporal depth changes using binocular depth maps, and introduces a confidence-weighted scoring and temporal consistency judgment strategy to accurately distinguish between dynamic and static feature points; finally, the above modules are integrated into the ORB-SLAM3 framework to construct a complete system that includes tracking, local mapping, and loop detection. Experiments on the Euroc and KITTI datasets demonstrate that compared to ORB-SLAM3, the proposed method reduces absolute trajectory error (ATE) by 8%~15% in indoor scenes and by 7%~35% in outdoor scenes, significantly enhancing robustness and positioning accuracy in dynamic environments.
文章引用:潘长征, 杨艺. 融合深度与语义信息的双目视觉SLAM[J]. 传感器技术与应用, 2025, 13(6): 827-837. https://doi.org/10.12677/jsta.2025.136081

参考文献

[1] 黄泽霞, 邵春莉. 深度学习下的视觉SLAM综述[J]. 机器人, 2023, 45(6): 756-768.
[2] Ai, Y., Rui, T., Yang, X., He, J., Fu, L., Li, J., et al. (2021) Visual SLAM in Dynamic Environments Based on Object Detection. Defence Technology, 17, 1712-1721. [Google Scholar] [CrossRef
[3] Liu, J., Meng, Z. and You, Z. (2020) A Robust Visual SLAM System in Dynamic Man-Made Environments. Science China Technological Sciences, 63, 1628-1636. [Google Scholar] [CrossRef
[4] Liu, C., Qin, J., Wang, S., Yu, L. and Wang, Y. (2022) Accurate RGB-D SLAM in Dynamic Environments Based on Dynamic Visual Feature Removal. Science China Information Sciences, 65, Article No. 202206. [Google Scholar] [CrossRef
[5] Yu, C., Liu, Z., Liu, X., Xie, F., Yang, Y., Wei, Q., et al. (2018) DS-SLAM: A Semantic Visual SLAM Towards Dynamic Environments. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, 1-5 October 2018, 1168-1174. [Google Scholar] [CrossRef
[6] Jonnalagadda, V.A. and Hashim, A.H. (2024) SegNet: A Segmented Deep Learning Based Convolutional Neural Network Approach for Drones Wildfire Detection. Remote Sensing Applications: Society and Environment, 34, Article ID: 101181.
[7] Bescos, B., Facil, J.M., Civera, J. and Neira, J. (2018) DynaSLAM: Tracking, Mapping, and Inpainting in Dynamic Scenes. IEEE Robotics and Automation Letters, 3, 4076-4083. [Google Scholar] [CrossRef
[8] Zhang, X., Wang, X. and Zhang, R. (2022) Dynamic Semantics SLAM Based on Improved Mask R-CNN. IEEE Access, 10, 126525-126535. [Google Scholar] [CrossRef
[9] Zhong, F., Wang, S., Zhang, Z., Chen, C. and Wang, Y. (2018) Detect-SLAM: Making Object Detection and SLAM Mutually Beneficial. 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe 12-15 March 2018, 1001-1010. [Google Scholar] [CrossRef
[10] Cheng, S., Sun, C., Zhang, S. and Zhang, D. (2023) SG-SLAM: A Real-Time RGB-D Visual SLAM toward Dynamic Scenes with Semantic and Geometric Information. IEEE Transactions on Instrumentation and Measurement, 72, 1-12. [Google Scholar] [CrossRef
[11] Liu, Y. and Miura, J. (2021) RDS-SLAM: Real-Time Dynamic SLAM Using Semantic Segmentation Methods. IEEE Access, 9, 23772-23785. [Google Scholar] [CrossRef
[12] Hu, Z., Zhao, J., Luo, Y. and Ou, J. (2022) Semantic SLAM Based on Improved DeepLabv3⁺ in Dynamic Scenarios. IEEE Access, 10, 21160-21168. [Google Scholar] [CrossRef
[13] Niu, J., Chen, Z., Zhang, T. and Zheng, S. (2024) Improved Visual SLAM Algorithm Based on Dynamic Scenes. Applied Sciences, 14, Article 10727. [Google Scholar] [CrossRef
[14] 于明洋, 徐海青, 张文焯, 等. 融合ASPP与双注意力机制的建筑物提取模型[J]. 航天返回与遥感, 2024, 45(1): 136-146.
[15] 夏晓华, 苏建功, 王耀耀, 等. 基于DeepLabv3+的轻量化路面裂缝检测模型[J]. 激光与光电子学进展, 2024, 61(8): 182-191.
[16] Howard, A., Sandler, M., Chen, B., Wang, W., Chen, L., Tan, M., et al. (2019) Searching for MobileNetV3. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, 27 October-2 November 2019, 1314-1324. [Google Scholar] [CrossRef
[17] Yao, Y. (2023) SE-CNN: Convolution Neural Network Acceleration via Symbolic Value Prediction. IEEE Journal on Emerging and Selected Topics in Circuits and Systems, 13, 73-85. [Google Scholar] [CrossRef
[18] Zhao, D., Zhang, W. and Wang, Y. (2024) Research on Personnel Image Segmentation Based on Mobilenetv2 H-Swish CBAM PSPNet in Search and Rescue Scenarios. Applied Sciences, 14, Article 10675. [Google Scholar] [CrossRef
[19] Luo, J., Fang, H., Shao, F., Zhong, Y. and Hua, X. (2021) Multi-Scale Traffic Vehicle Detection Based on Faster R-CNN with NAS Optimization and Feature Enrichment. Defence Technology, 17, 1542-1554. [Google Scholar] [CrossRef
[20] Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W. and Hu, Q. (2020) ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 11531-11539. [Google Scholar] [CrossRef
[21] 王梦瑶, 宋薇. 动态场景下基于自适应语义分割的RGB-DSLAM算法[J]. 机器人, 2023, 45(1): 16-27.