融合胶囊网络与特征金字塔网络的YOLOv12小麦病害检测
Wheat Disease Detection Based on YOLOv12 with Capsule Networks and FPN
DOI: 10.12677/airr.2026.155109, PDF,   
作者: 幸南霖*, 余 洪:重庆三峡科技大学计算机科学与工程学院,重庆;林 训*, 唐欣怡, 陈 杭#:重庆旅游职业学院智慧旅游学院,重庆
关键词: 小麦病害检测YOLOv12胶囊网络特征金字塔网络小目标检测Wheat Disease Detection YOLOv12 Capsule Networks FPN Small Object Detection
摘要: 小麦病害严重威胁全球粮食安全,造成巨大的粮食减产。传统人工病害诊断方式效率低下、判定主观性强,无法适用于大田规模化检测场景。基于YOLO系列的目标检测模型虽检测速度较快,但该类算法缺少对目标空间结构信息的建模能力,在复杂田间背景下针对小尺寸病斑、被遮挡病斑的识别效果较差。本文将胶囊网络(CapsNet)与特征金字塔网络(FPN)嵌入YOLOv12算法,构建一种改进型小麦病害检测模型。研究搭建包含7245张小麦病害图像的数据集并完成数据增广,按照7:2:1的比例划分训练集、验证集与测试集。引入胶囊网络强化模型对空间特征的表征能力,依托特征金字塔网络优化多尺度特征融合效果。实验结果表明,本文所提模型精确率可达92.5%,召回率90.2%,F1值91.3%,交并比阈值0.5下平均精度均值(mAP@0.5)为93.8%,mAP@0.5:0.95指标为69.7%。消融实验验证了各改进模块的有效性;相较于原生YOLOv12,本模型平均精度提升4.8%,对微小病斑具有更优异的检测性能。经统计学t检验分析,模型性能提升具有统计学显著性(p < 0.01)。该方法可为精准农业领域小麦病害智能化识别检测提供可行有效的技术方案。
Abstract: Wheat diseases severely threaten global grain security and cause substantial yield losses. Traditional manual diagnosis is inefficient, subjective, and unsuitable for large-scale fields. Although YOLO-based detectors achieve high speed, they lack spatial structural modeling and perform poorly on small, occluded lesions under complex backgrounds. This paper proposes an improved model by integrating Capsule Networks (CapsNet) and Feature Pyramid Network (FPN) into YOLOv12. A dataset of 7245 wheat disease images is constructed and augmented, then split into training/validation/test sets at 7:2:1. CapsNet enhances spatial feature representation, while FPN strengthens multi-scale fusion. Experiments show the proposed method achieves 92.5% Precision, 90.2% Recall, 91.3% F1-score, 93.8% mAP@0.5, and 69.7% mAP@0.5:0.95. Ablation studies verify the effectiveness of each module. Compared with YOLOv12, the model improves APs by 4.8%, demonstrating superior performance on tiny lesions. Statistical t-tests confirm significant improvements (p < 0.01). This method provides an effective solution for intelligent wheat disease detection in precision agriculture.
文章引用:幸南霖, 林训, 余洪, 唐欣怡, 陈杭. 融合胶囊网络与特征金字塔网络的YOLOv12小麦病害检测[J]. 人工智能与机器人研究, 2026, 15(5): 1203-1212. https://doi.org/10.12677/airr.2026.155109

参考文献

[1] Mohanty, S.P., Hughes, D.P. and Salathé, M. (2016) Using Deep Learning for Image-Based Plant Disease Detection. Frontiers in Plant Science, 7, Article 1419.
https://doi.org/10.3389/fpls.2016.01419
[2] Ferentinos, K.P. (2018) Deep Learning Models for Plant Disease Detection and Diagnosis. Computers and Electronics in Agriculture, 145, 311-318.
https://doi.org/10.1016/j.compag.2018.01.009
[3] Too, E.C., Yujian, L., Njuki, S. and Yingchun, L. (2019) A Comparative Study of Fine-Tuning Deep Learning Models for Plant Disease Identification. Computers and Electronics in Agriculture, 161, 272-279.
https://doi.org/10.1016/j.compag.2018.03.032
[4] Redmon, J., Divvala, S., Girshick, R. and Farhadi, A. (2016) You Only Look Once: Unified, Real-Time Object Detection. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, 27-30 June 2016, 779-788.
https://doi.org/10.1109/cvpr.2016.91
[5] Redmon, J. and Farhadi, A. (2018) YOLOv3: An Incremental Improvement. arXiv: 1804.02767.
[6] Bochkovskiy, A., Wang, C.Y. and Liao, H.Y.M. (2020) YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv: 2004.10930.
[7] Wang, C., Bochkovskiy, A. and Liao, H.M. (2023) YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 7464-7475.
https://doi.org/10.1109/cvpr52729.2023.00721
[8] Sabour, S., Frosst, N. and Hinton, G.E. (2017) Dynamic Routing between Capsules. arXiv: 1710.09829.
[9] Hinton, G.E., Sabour, S. and Frosst, N. (2018) Matrix Capsules with EM Routing. ICLR 2018, Vancouver, 30 April-3 May 2018, 1-15.
[10] Lin, T., Dollar, P., Girshick, R., He, K., Hariharan, B. and Belongie, S. (2017) Feature Pyramid Networks for Object Detection. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, 21-26 July 2017, 936-944.
https://doi.org/10.1109/cvpr.2017.106
[11] Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J. and Zisserman, A. (2009) The Pascal Visual Object Classes (VOC) Challenge. International Journal of Computer Vision, 88, 303-338.
https://doi.org/10.1007/s11263-009-0275-4
[12] Tian, Y., Ye, Q. and Doermann, D. (2025) YOLOv12: Attention-Centric Real-Time Object Detectors. Advances in Neural Information Processing Systems 38, San Diego, 2-7 December 2025, 87151-87175.
https://doi.org/10.52202/085713-2627
[13] Tian, Y., Ye, Q. and Doermann, D. (2025) sunsmarterjie/yolov12.
https://github.com/sunsmarterjie/yolov12
[14] Ren, S., He, K., Girshick, R. and Sun, J. (2017) Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39, 1137-1149.
https://doi.org/10.1109/tpami.2016.2577031
[15] Zhao, Z., Zheng, P., Xu, S. and Wu, X. (2019) Object Detection with Deep Learning: A Review. IEEE Transactions on Neural Networks and Learning Systems, 30, 3212-3232.
https://doi.org/10.1109/tnnls.2018.2876865