基于结构感知花瓣拓扑图的细粒度牡丹识别
Fine-Grained Peony Recognition Based on a Structure-Aware Petal Topology Graph
DOI: 10.12677/jisp.2026.153036, PDF,   
作者: 陈寒睿*#, 鞠奥宇*, 郑宏艺, 李军保, 刘扬帆:洛阳理工学院电子信息学院,河南 洛阳;张 玉:洛阳理工学院材料学院,河南 洛阳;李 真:洛阳理工学院人工智能学院,河南 洛阳
关键词: 牡丹识别花瓣拓扑实例分割图注意力网络结构感知分类Peony Recognition Petal Topology Instance Segmentation Graph Attention Network Structure-Aware Classification
摘要: 目的:针对牡丹图像中重瓣花冠遮挡、花瓣边界粘连和局部纹理相近等问题,构建基于花瓣拓扑关系的识别模型PTRGNet。方法:训练阶段采用SAM候选掩码与人工校正获得实例级花瓣标注;测试阶段由训练后的实例分割器自动输出花瓣候选。模型以花瓣实例为节点,融合空间形态特征和ROI语义特征,并将距离、方向角、重叠率、外接框IoU和面积差作为边属性输入边感知GAT,再与视觉分支门控融合。结果:在Peony-8上,PTRGNet取得67.8% Top-1准确率和67.1% F1分数;在Oxford Flowers-102上,弱伪拓扑设置取得89.8% Top-1和98.6% Top-5。结论:花瓣拓扑可为牡丹识别提供颜色、纹理之外的结构证据,但模型性能仍受实例分割质量影响。
Abstract: Objective: Peony recognition is often disturbed by dense petals, weak boundaries and similar local textures. This work builds a Petal Topology Relation Graph Network (PTRGNet) to introduce structural cues into fine-grained peony classification. Methods: SAM is first used as an annotation aid to produce candidate petal masks, which are then corrected manually for instance-level supervision. At test time, petal masks are generated by the trained instance segmenter without manual input. PTRGNet represents each petal as a graph node, combines spatial morphology with ROI semantic features, and describes pairwise relations using distance, direction, mask overlap, bounding-box IoU and relative area difference. These edge attributes are fed into an edge-aware graph attention module and fused with the visual branch through a gated head. Results: On Peony-8, PTRGNet reached 67.8% Top-1 accuracy, 66.9% precision, 67.3% recall and 67.1% F1 score under predicted masks, exceeding EfficientNet-B0 by 8.1 percentage points. Human-refined masks gave 70.2% accuracy, while zero-shot SAM masks gave 63.5%, showing the effect of mask noise. On Oxford Flowers-102, the weak pseudo-topology setting achieved 89.8% Top-1 and 98.6% Top-5 accuracy. Conclusion: Petal topology supplies useful structural evidence for peony recognition, although stronger validation still depends on reliable automatic masks and larger instance-level flower datasets.
文章引用:陈寒睿, 鞠奥宇, 郑宏艺, 李军保, 张玉, 李真, 刘扬帆. 基于结构感知花瓣拓扑图的细粒度牡丹识别[J]. 图像与信号处理, 2026, 15(3): 401-414. https://doi.org/10.12677/jisp.2026.153036

参考文献

[1] 罗建豪, 吴建鑫. 基于深度卷积特征的细粒度图像分类研究综述[J]. 自动化学报, 2017, 43(8): 1306-1318.
[2] Nilsback, M. and Zisserman, A. (2008) Automated Flower Classification over a Large Number of Classes. 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, Bhubaneswar, 16-19 December 2008, 722-729.
https://doi.org/10.1109/icvgip.2008.47
[3] 李嘉珏. 中国牡丹与芍药[M]. 北京: 中国林业出版社, 1999.
[4] 陈少真, 叶武剑, 刘怡俊. 基于知识蒸馏与改进ViT网络的花卉图像细粒度分类[J]. 光电子∙激光, 2024, 35(1): 29-40.
[5] 廖宁, 曹敏, 严骏驰. 视觉提示学习综述[J]. 计算机学报, 2024, 47(4): 790-820.
[6] Fu, J., Zheng, H. and Mei, T. (2017) Look Closer to See Better: Recurrent Attention Convolutional Neural Network for Fine-Grained Image Recognition. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, 21-26 July 2017, 4476-4484.
https://doi.org/10.1109/cvpr.2017.476
[7] He, J., Chen, J., Liu, S., Kortylewski, A., Yang, C., Bai, Y., et al. (2022) TransFG: A Transformer Architecture for Fine-Grained Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 36, 852-860.
https://doi.org/10.1609/aaai.v36i1.19967
[8] 项伟康, 周全, 崔景程, 莫智懿, 吴晓富, 欧卫华, 王井东, 刘文予. 基于深度学习的弱监督语义分割方法综述[J]. 中国图象图形学报, 2024, 29(5): 1146-1168.
[9] He, K., Gkioxari, G., Dollar, P. and Girshick, R. (2017) Mask R-CNN. 2017 IEEE International Conference on Computer Vision (ICCV), Venice, 22-29 October 2017, 2980-2988.
https://doi.org/10.1109/iccv.2017.322
[10] Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al. (2023) Segment Anything. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, 1-6 October 2023, 3992-4003.
https://doi.org/10.1109/iccv51070.2023.00371
[11] 徐冰冰, 岑科廷, 黄俊杰, 沈华伟, 程学旗. 图卷积神经网络综述[J]. 计算机学报, 2020, 43(5): 755-780.
[12] 马帅, 刘建伟, 左信. 图神经网络综述[J]. 计算机研究与发展, 2022, 59(1): 47-80.
[13] Jing, D., He, X., Luo, Y., Fei, N., Yang, G., Wei, W., et al. (2024) FineCLIP: Self-Distilled Region-Based CLIP for Better Fine-Grained Understanding. Advances in Neural Information Processing Systems 37, Vancouver, 10-15 December 2024, 27896-27918.
https://doi.org/10.52202/079017-0875
[14] Wang, Z., Zhang, Z., Chawla, N., Zhang, C. and Ye, Y. (2024) GFT: Graph Foundation Model with Transferable Tree Vocabulary. Advances in Neural Information Processing Systems 37, Vancouver, 10-15 December 2024, 107403-107443.
https://doi.org/10.52202/079017-3412
[15] Xu, D.F., Zhu, Y.K., Choy, C.B. and Li, F.F. (2017) Scene Graph Generation by Iterative Message Passing. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, 21-26 July 2017, 3097-3106.
https://doi.org/10.1109/cvpr.2017.330