基于跨摄像头的车辆在线跟踪关键技术研究
DSESHADE: Research on Key Technologies for Vehicle Online Tracking Based on Cross-Camera Association
摘要: 多摄像头多目标车辆跟踪(Multi-Target Multi-Camera Vehicle Tracking, MTMCT)由于车辆外观变化、视角差异以及跨摄像头域分布偏移等问题,仍然面临较大挑战。车辆重识别(Vehicle re-identification, ReID)作为跨摄像头轨迹关联的关键技术,在现有数据集约束下,现有方法主要依赖外观特征建模,导致模型鲁棒性与跨域泛化能力受限。针对上述问题,本文构建了大规模多维标注综合数据集VeCamX,通过融合VeRi-776丰富的语义属性信息与AIC22复杂的时空轨迹特征,实现车辆身份、属性、轨迹及姿态关键点等多层次标注统一,为在线车辆跟踪 提供综合性评测基准。 在此基础上,提出一种统一优化框架,融合姿态感知模块(Pose-Aware Module, PAM)、无数据辅助模块(Metadata-Assisted Module, MAM)以及搜索剪枝优化策 略(Search-and-Pruning, SnP)。其中,PAM用于学习车辆目标的视角不变结构特征;MAM 整合车辆类型、 晶牌、 颜色以及时空轨迹等无数据信息;SnP通过构建均衡且高判别性的训练样本集提升模型泛化能力。 在VeRi-776、 AIC22及VeCamX数据集上的实验结果表明,所提出框 架能够稳定提升IDF1 和MOTA 等评价指标,相较现有先进方法表现出更强的鲁棒性与跨场景泛化能力。
Abstract: Multi-Target Multi-Camera vehicle Tracking (MTMCT) remains challenging due to severe appearance variations, viewpoint shifts, and domain discrepancies among cameras. Notably, vehicle re-identification (ReID) plays a crucial role in trajectory association for MTMCT, yet constrained by the limitations of existing datasets, cur- rent ReID-based approaches rely mainly on appearance cues, which compromises their robustness and cross-domain generalization. To address these issues, we in- troduce a large-scale, multi-dimensionally annotated dataset named VeCamX, which integrates the semantic richness of VeRi-776 and the spatio-temporal complexity of AIC22. Specifically, VeCamX unifies annotations for vehicle identities, attributes, trajectories, and poses keypoints, thereby providing a comprehensive benchmark for online vehicle tracking.Building upon VeCamX’s multi-dimensional annotations, we propose a unified optimization framework that combines Pose-Aware Module (PAM), Metadata-Assisted Module (MAM), and Search-and-Pruning optimization strategy (SnP). Specifically, PAM captures viewpoint-invariant structural features of vehicle targets; MAM acquires target-specific metadata including vehicle type, brand, color, and spatio-temporal trajectories; and SnP supplies balanced and highly discrimina- tive data inputs for model training, thereby enhancing the model’s generalization performance. Experiments on VeRi-776, AIC22, and VeCamX show that the pro- posed framework consistently improves IDF1 and MOTA over state-of-the-art meth- ods, achieving strong robustness and generalization in complex real-world scenarios.
文章引用:黄巍鑫. 基于跨摄像头的车辆在线跟踪关键技术研究[J]. 计算机科学与应用, 2026, 16(7): 226-241. https://doi.org/10.12677/CSA.2026.167254

参考文献

[1] Veres, M. and Moussa, M. (2020) Deep Learning for Intelligent Transportation Systems: A Survey of Emerging Trends. IEEE Transactions on Intelligent Transportation Systems, 21, 3152-3168. [Google Scholar] [CrossRef
[2] Shim, K., Ko, K., Hwang, J., Jang, H. and Kim, C. (2023) Fast Online Multi-Target Multi- Camera Tracking for Vehicles. Applied Intelligence, 53, 28994-29004. [Google Scholar] [CrossRef
[3] Yang, H., Cai, J., Liu, C., Ke, R. and Wang, Y. (2023) Cooperative Multi-Camera Ve- hicle Tracking and Traffic Surveillance with Edge Artificial Intelligence and Representation Learning. Transportation Research Part C: Emerging Technologies, 148, Article ID: 103982. [Google Scholar] [CrossRef
[4] Fei, L. and Han, B. (2023) Multi-Object Multi-Camera Tracking Based on Deep Learning for Intelligent Transportation: A Review. Sensors, 23, Article 3852. [Google Scholar] [CrossRef
[5] Wang, H., Hou, J. and Chen, N. (2019) A Survey of Vehicle Re-Identification Based on Deep Learning. IEEE Access, 7, 172443-172469. [Google Scholar] [CrossRef
[6] Lou, Y., Bai, Y., Liu, J., Wang, S. and Duan, L. (2019) VERI-Wild: A Large Dataset and a New Method for Vehicle Re-Identification in the Wild. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, 15-20 June 2019, 3230-3238. [Google Scholar] [CrossRef
[7] Li, F., Wang, Z., Nie, D., Zhang, S., Jiang, X., Zhao, X., et al. (2022) Multi-Camera Vehicle Tracking System for AI City Challenge 2022. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, 19-20 June 2022, 3264-3272. [Google Scholar] [CrossRef
[8] Zhang, F., Zhang, L., Zhang, H. and Ma, Y. (2023) Image-To-Image Domain Adaptation for Vehicle Re-Identification. Multimedia Tools and Applications, 82, 40559-40584. [Google Scholar] [CrossRef
[9] Wei, R., Gu, J., He, S. and Jiang, W. (2023) Transformer-Based Domain-Specific Represen- tation for Unsupervised Domain Adaptive Vehicle Re-Identification. IEEE Transactions on Intelligent Transportation Systems, 24, 2935-2946. [Google Scholar] [CrossRef
[10] Li, J., Zhang, S., Tian, Q., Wang, M. and Gao, W. (2022) Pose-Guided Representation Learn- ing for Person Re-identification. IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 44, 622-635. [Google Scholar] [CrossRef
[11] Hsu, H., Cai, J., Wang, Y., Hwang, J. and Kim, K. (2021) Multi-Target Multi-Camera Tracking of Vehicles Using Metadata-Aided Re-Id and Trajectory-Based Camera Link Model. IEEE Transactions on Image Processing, 30, 5198-5210. [Google Scholar] [CrossRef
[12] Jain, E., Nandy, T., Aggarwal, G., Tendulkar, A., Iyer, R. and De, A. (2023) Efficient Data Subset Selection to Generalize Training across Models: Transductive and Inductive Networks. Advances in Neural Information Processing Systems 36, New Orleans, 10-16 December 2023, 4716-4740. [Google Scholar] [CrossRef
[13] Xue, C., Deng, Z., Wang, S., Hu, E., Zhang, Y., Yang, W., et al. (2024) GLSFF: Global-Local Specific Feature Fusion for Cross-Modality Pedestrian Re-Identification. Computer Communi- cations, 215, 157-168. [Google Scholar] [CrossRef
[14] Zhai, Y., Zeng, Y., Huang, Z., Qin, Z., Jin, X. and Cao, D. (2024) Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification. Proceedings of the AAAI Conference on Artificial Intelligence, 38, 6979-6987. [Google Scholar] [CrossRef
[15] Nguyen, T.T., Nguyen, H.H., Sartipi, M. and Fisichella, M. (2024) Multi-Vehicle Multi-Camera Tracking with Graph-Based Tracklet Features. IEEE Transactions on Multimedia, 26, 972-983. [Google Scholar] [CrossRef
[16] Chen, Y.C., Jing, L.L., Vahdani, E., Zhang, L., He, M.Y. and Tian, Y.L. (2019) Multi-Camera Vehicle Tracking and Re-Identification on AI City Challenge 2019. CVPR Workshops 2019, Long Besch, 16-20 June 2019, 324-332.
[17] Yoon, Y., Song, Y., Yoon, K. and Jeon, M. (2018) Online Multi-Object Tracking Using Selective Deep Appearance Matching. 2018 IEEE International Conference on Consumer Electronics—Asia (ICCE-Asia), JeJu, 24-26 June 2018, 206-212. [Google Scholar] [CrossRef
[18] Nguyen, T.T., Nguyen, H.H., Sartipi, M. and Fisichella, M. (2024) Lammon: Language Model Combined Graph Neural Network for Multi-Target Multi-Camera Tracking in Online Scenarios.
[19] Machine Learning, 113, 6811-6837.[CrossRef
[20] Nguyen, T.T., Nguyen, H.H., Sartipi, M. and Fisichella, M. (2023) Real-Time Multi-Vehicle Multi-Camera Tracking with Graph-Based Tracklet Features. Transportation Research Record: Journal of the Transportation Research Board, 2678, 296-308. [Google Scholar] [CrossRef
[21] Ayala-Acevedo, A., Devgun, A., Zahir, S. and Askary, S. (2019) Vehicle Re-Identification: Pushing the Limits of Re-Identification. CVPR Workshops 2019, Long Beach, 16-20 June 2019, 291-296.
[22] Wang, C., Bochkovskiy, A. and Liao, H.M. (2023) YOLOv7: Trainable Bag-Of-Freebies Set- s New State-Of-The-Art for Real-Time Object Detectors. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 7464-7475. [Google Scholar] [CrossRef
[23] Zhang, Y., Sun, P., Jiang, Y., Yu, D., Weng, F., Yuan, Z., et al. (2022) ByteTrack: Multi- Object Tracking by Associating Every Detection Box. In: Avidan, S., Brostow, G., Cisse, M., Farinella, G.M. and Hassner, T., Eds., Computer Vision—ECCV 2022, Springer, 1-21. 1 [Google Scholar] [CrossRef
[24] Sun, K., Xiao, B., Liu, D. and Wang, J. (2019) Deep High-Resolution Representation Learning for Human Pose Estimation. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Kushtia, 23-24 October 2025, 1-6. [Google Scholar] [CrossRef