面向无标签数据的三维场景语义理解方法研究
Research on Semantic Understanding Methods for 3D Scenes with Unlabeled Data
DOI: 10.12677/airr.2026.154099, PDF,    科研立项经费支持
作者: 高 丽:沈阳城市学院生命与健康管理学院,辽宁 沈阳;王一可, 尹竞瑶:沈阳城市学院智能与工程学院,辽宁 沈阳
关键词: 三维场景理解无标签学习三维高斯溅射面向对象表达自监督学习3D Scene Understanding Unsupervised Learning 3D Gaussian Splatting Object-Oriented Representation Self-Supervised Learning
摘要: 三维场景理解是计算机视觉领域的核心课题,其目标是从二维观测数据中恢复场景的三维结构并赋予语义含义。现有方法通常依赖大量带标签数据,且多采用面向任务的场景表达方式,难以兼顾表达能力的全面性与学习效率。针对上述问题,本文提出一种面向对象的三维场景语义理解框架,旨在从无标签数据中学习以对象为中心的连续场景表达。该方法以三维高斯溅射(3D Gaussian Splatting, 3DGS)为核心表示基元,设计了一种融合自监督与对抗学习的复合训练策略:一方面利用几何与物理一致性提供自监督信号,另一方面通过多模态先验知识引导场景表达的规则化。实验结果表明,该方法在三维重建质量和语义理解能力方面均取得了良好效果,验证了面向对象表达与无标签学习的有效结合。本研究为实现低标注依赖、高泛化能力的三维场景理解提供了可行路径。
Abstract: Three-dimensional scene understanding is a fundamental task in computer vision, aiming to recover the 3D structure of a scene from 2D observations while assigning semantic meaning to it. Existing approaches typically rely heavily on large amounts of labeled data and often adopt task-specific scene representations, which struggle to balance representational expressiveness with learning efficiency. To address these challenges, this paper proposes an object-oriented framework for 3D scene semantic understanding, designed to learn continuous, object-centric scene representations from unlabeled data. The proposed method employs 3D Gaussian Splatting (3DGS) as the core representation primitive and introduces a hybrid training strategy that integrates self-supervised learning with adversarial learning. Specifically, geometric and physical consistency constraints are leveraged to provide self-supervised signals, while multimodal prior knowledge is incorporated to regularize the scene representation. Experimental results demonstrate that the proposed method achieves promising performance in both 3D reconstruction quality and semantic understanding capability, validating the effective combination of object-oriented representation and label-free learning. This study offers a viable pathway toward 3D scene understanding with reduced annotation dependence and enhanced generalization ability.
文章引用:高丽, 王一可, 尹竞瑶. 面向无标签数据的三维场景语义理解方法研究[J]. 人工智能与机器人研究, 2026, 15(4): 1097-1104. https://doi.org/10.12677/airr.2026.154099

参考文献

[1] Gao, Z., Yi, R., Huang, Y., Chen, W., Zhu, C. and Xu, K. (2025) Self-Supervised Learning of Hybrid Part-Aware 3D Representations of 2D Gaussians and Superquadrics. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-25 October 2025, 9649-9659. [Google Scholar] [CrossRef
[2] 连振晗, 王章野, 牛炯, 程乐超. 基于三维高斯泼溅的动态场景重建研究综述[J]. 计算机辅助设计与图形学学报, 2026, 38(1): 61-77.
[3] 从“像素”到“三维”: 3DGS如何革新数字世界的沉浸体验? [EB/OL]. 科普中国.
https://cloud.kepuchina.cn/newSearch/imgText?from=1&id=7358183627336155136&is_self=2, 2025-08-22.
[4] Kerbl, B., Kopanas, G., Leimkuehler, T. and Drettakis, G. (2023) 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics, 42, 1-14. [Google Scholar] [CrossRef
[5] Luan, F., Sun, S., Zhang, H., Yin, Y., Wang, K. and Yang, J. (2025) 3D Gaussian Splatting Technologies and Extensions: A Review. Neurocomputing, 658, Article 131629. [Google Scholar] [CrossRef
[6] Yamada, R., Ide, K., Fukuhara, Y., et al. (2024) 3D Sans 3D Scans: Scalable Pre-Training from Video-Generated Point Clouds. arXiv:2512.23042.
[7] Zanjani, F.G., Cai, H., Ackermann, H., Mirvakhabova, L. and Porikli, F. (2025) Planar Gaussian Splatting. 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, 26 February 2025-6 March 2025, 8905-8914. [Google Scholar] [CrossRef
[8] Li, Y., Ma, Q., Yang, R., Li, H., Ma, M., Ren, B., et al. (2025) Scenesplat: Gaussian Splatting-Based Scene Understanding with Vision-Language Pretraining. 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, 19-25 October 2025, 4961-4972. [Google Scholar] [CrossRef
[9] Huang, S., Xie, Y., Zhu, S. and Zhu, Y. (2021) Spatio-Temporal Self-Supervised Representation Learning for 3D Point Clouds. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 6515-6525. [Google Scholar] [CrossRef
[10] Zhang, W., Zhang, L., Hu, P., Ma, L., Zhuge, Y. and Lu, H. (2025) Bootstraping Clustering of Gaussians for View-Consistent 3D Scene Understanding. Proceedings of the AAAI Conference on Artificial Intelligence, 39, 10166-10175. [Google Scholar] [CrossRef
[11] Nguyen-Phuoc, T., Richardt, C., Mai, L., et al. (2020) BlockGAN: Learning 3D Object-Aware Scene Representations from Unlabelled Images. arXiv:2002.08988.
[12] InteriorGS——群核科技推出的高质量3D高斯语义数据集[EB/OL]. 2025-07-20.
https://ai-bot.cn/interiorgs/, 2026-06-11.
[13] Kim, S., Cheng, Y., Kong, X., Kelly, P.H.J. and Davison, A.J. (2026) MLP Splatting: Object-Centric Neural Fields. arXiv:2606.03877.
[14] Hengye, L., Yanli, L., Hong, L., Xia, Y. and Guanyu, X. (2025) Multi-View Intrinsic Decomposition of Indoor Scenes under a 3D Gaussian Splatting Framework. Journal of Image and Graphics, 30, 2514-2527. [Google Scholar] [CrossRef