改进的YOLOv11s无人机手势识别算法
Improved YOLOv11s Algorithm for UAV Gesture Recognition
摘要: 针对标准YOLOv11算法在直接应用于无人机手势识别时,面临手部特征提取不充分、多尺度手势适应能力不足以及边界框定位精度有限等问题,导致检测效果难以满足实际需求。为此,本文从特征提取、多尺度融合及损失函数三个方向对YOLOv11进行改进,提出适用于无人机手势识别任务的改进算法。首先,针对手部特征提取能力不足的问题,在Backbone阶段采用C2PSA_EDFFN模块替换原有的C2f模块。该模块通过增强特征提取过程中的上下文信息交互,有效提升了网络对手部关键特征的捕获能力,改善了对手势区域的鉴别效果。其次,针对不同尺度手势(如远距离小目标手势与近距离大目标手势)的检测适应性问题,在Neck阶段引入CGAFusion模块优化多尺度特征融合策略。该模块能够自适应整合浅层细节信息与深层语义信息,提升了模型对不同尺度手势的检测能力。最后,针对手势边界框定位精度不足的问题,采用PIoU2损失函数替代原有的CIoU损失函数。该损失函数能够更精确地评估预测框与真实框之间的重叠程度,提高了手势目标边界框的回归精度。
Abstract: When directly applied to UAV-based hand gesture recognition, the standard YOLOv11 algorithm suffers from insufficient hand feature extraction, limited adaptability to multi-scale gestures, and suboptimal bounding box localization accuracy, making it difficult to satisfy practical application requirements. To address these limitations, this paper proposes an improved YOLOv11-based algorithm tailored for UAV hand gesture recognition by enhancing feature extraction, multi-scale feature fusion, and loss function design. Specifically, to improve hand feature representation, the C2PSA_EDFFN module is introduced in the backbone to replace the original C2f module, which enhances contextual information interaction during feature extraction and strengthens the network’s ability to capture critical hand features. To better handle gestures at different scales, a CGAFusion module is incorporated into the neck to optimize multi-scale feature fusion by adaptively integrating shallow detailed information with deep semantic features, thereby improving detection performance across varying gesture scales. Furthermore, to enhance bounding box localization accuracy, the PIoU2 loss function is employed to replace the original CIoU loss, enabling a more precise evaluation of the overlap between predicted and ground-truth bounding boxes and improving regression accuracy for gesture localization.
文章引用:迟晨, 魏立臻, 赵菁菁, 朱以墨, 张丽艳. 改进的YOLOv11s无人机手势识别算法[J]. 人工智能与机器人研究, 2026, 15(4): 1085-1096. https://doi.org/10.12677/airr.2026.154098

参考文献

[1] Zhang, B., Zhang, H., Zhen, T., Ji, B., Xie, L., Yan, Y., et al. (2024) A Two-Stage Real-Time Gesture Recognition Framework for UAV Control. IEEE Sensors Journal, 24, 24770-24782.
https://doi.org/10.1109/jsen.2024.3413787
[2] Fang, W., Lai, G., Yu, Y., Shi, C., Meng, X., Sun, J., et al. (2025) Real-Time Human-Drone Interaction via Active Multimodal Gesture Recognition under Limited Field of View in Indoor Environments. IEEE Robotics and Automation Letters, 10, 11705-11712.
https://doi.org/10.1109/lra.2025.3615031
[3] Li, S., Wang, Z., Liang, J. and Wang, Y. (2025) Small Object Detection in UAV Scenarios Based on YOLOv5. Computer Modeling in Engineering & Sciences, 145, 3993-4011.
https://doi.org/10.32604/cmes.2025.073896
[4] Guo, N., Ding, Y. and Meng, H. (2025) YOLOv8-GR: Real Time Gesture Recognition by Improving of YOLOv8 Algorithm. Complex & Intelligent Systems, 11, Article No. 479.
https://doi.org/10.1007/s40747-025-02107-0
[5] Liang, X., Xiang, J., Qin, S., Xiao, Y., Chen, L., Zou, D., et al. (2025) Small Target Detection Algorithm Based on Sahi-Improved-Yolov8 for UAV Imagery: A Case Study of Tree Pit Detection. Smart Agricultural Technology, 12, Article 101181.
https://doi.org/10.1016/j.atech.2025.101181
[6] Zhang, B., Li, J., Wang, H., et al. (2025) Object Detection Model of Vehicle-Road Cooperative Autonomous Driving Based on Improved YOLO11 Algorithm. PLOS ONE, 20, e0329628.
[7] Wang, C., Zhang, K., Ma, J., Chen, Z. and Yu, Y. (2025) Improved YOLOv11n-Based Small Object Detection Algorithm from UAV Perspective. 2025 4th International Symposium on Computer Applications and Information Technology (ISCAIT), Xi’an, China, 21-23 March 2025, 900-908.
https://doi.org/10.1109/iscait64916.2025.11010315
[8] Zheng, L., Wang, Z., Chen, Y., et al. (2025) Efficient Discriminative Frequency Domain-Based Feedforward Network for Image Restoration. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 11-15 June 2025, 25641-25650.
[9] Chen, J., Liu, Y., Zhang, S., et al. (2024) Powerful-IoU and PIoU v2: Advanced Boundary Box Regression Losses for Object Detection. Neural Networks, 179, Article ID: 106542.
[10] Zhao, Q. and Zhu, J. (2025) An Improved YOLOv11 Architecture with Multi-Scale Attention and Spatial Fusion for Fine-Grained Residual Detection. Results in Engineering, 27, Article 107061.
https://doi.org/10.1016/j.rineng.2025.107061
[11] Kong, L., Dong, J., Ge, J., Li, M. and Pan, J. (2023) Efficient Frequency Domain-Based Transformers for High-Quality Image Deblurring. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 5886-5895.
https://doi.org/10.1109/cvpr52729.2023.00570
[12] Chen, Z., He, Z. and Lu, Z.M. (2024) DEA-Net: Single Image Dehazing Based on Detail-Enhanced Convolution and Content-Guided Attention. IEEE Transactions on Image Processing, 33, 1002-1015.
https://doi.org/10.1109/tip.2024.3354108
[13] Li, X., Wang, Z., Chen, Y., et al. (2025) A Drone Detection Network Based on Multi-Scale Feature Fusion and Content-Guided Attention Mechanism. 2025 8th International Conference on Artificial Intelligence and Big Data (ICAIBD), Chengdu, 23-26 May 2025, 1-6.
[14] Liu, C., Wang, K., Li, Q., Zhao, F., Zhao, K. and Ma, H. (2024) Powerful-IoU: More Straightforward and Faster Bounding Box Regression Loss with a Nonmonotonic Focusing Mechanism. Neural Networks, 170, 276-284.
https://doi.org/10.1016/j.neunet.2023.11.041
[15] Kapitanov, A., Kvanchiani, K., Nagaev, A., et al. (2024) HaGRID—HAnd Gesture Recognition Image Dataset. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, 3-8 January 2024, 4572-4581.