基于空–频域联合表征的YOLOv11改进模型及初期火灾烟雾检测应用
YOLOv11 Improved Model Based on Joint Spatial-Frequency Domain Representation and Its Application in Early Fire Smoke Detection
摘要: 针对复杂环境下初期火灾烟雾小目标、动态性检测难题,提出一种基于改进YOLOv11的高精度轻量化火灾烟雾检测算法。将空–频域聚合Mamba引入YOLOv11架构,构建MixFreSBlock模块替代原C3K2模块:通过MSPLCK模块在空域构建多尺度大感受野,同时利用MFCA模块在频域实现噪声压缩与纹理强化,二者拼接形成互补的空–频联合表征;采用SS2D模块沿四维扫描路径对联合特征进行线性复杂度全局建模,突破卷积局部性限制;最后嵌入坐标注意力机制,沿水平–垂直方向重校准特征,精准定位火焰、烟雾细长结构。实验结果表明,该模型在自建与开源数据集上mAP50达94.7%,较原YOLOv11提升0.9%,参数量仅3.1 M,计算量仅8.0 G,在精度与轻量化之间实现更优平衡。本研究为实时火灾预警系统提供了高性能、低计算成本的解决方案,可广泛应用于安防监控、智慧城市等多场景,具有重要的实际应用价值。
Abstract: To address the challenges of detecting small and dynamic early fire smoke targets in complex environments, a high-precision lightweight fire smoke detection algorithm based on an improved YOLOv11 is proposed. The spatial-frequency domain aggregation Mamba is introduced into the YOLOv11 architecture, and the MixFreSBlock module is designed to replace the original C3K2 module: the MSPLCK module constructs a multi-scale large receptive field in the spatial domain, while the MFCA module achieves noise compression and texture enhancement in the frequency domain, and the two are concatenated to form complementary spatial-frequency joint representations; the SS2D module performs linear complexity global modelling of joint features along a four-dimensional scanning path, overcoming the locality limitation of convolution; finally, a coordinate attention mechanism is embedded to recalibrate features along horizontal and vertical directions, accurately locating flame and slender smoke structures. Experimental results show that the model achieves an mAP50 of 94.7% on both custom and open-source datasets, an increase of 0.9% compared to the original YOLOv11, with only 3.1 M parameters and 8.0 G computations, achieving a better balance between accuracy and lightweight design. This study provides a high-performance, low-computation solution for real-time fire warning systems and can be widely applied in security monitoring, smart cities, and other scenarios, with significant practical value.
文章引用:陈漓, 崔燕, 邓乾涛. 基于空–频域联合表征的YOLOv11改进模型及初期火灾烟雾检测应用[J]. 安防技术, 2026, 14(2): 57-72. https://doi.org/10.12677/jsst.2026.142006

参考文献

[1] Gajendiran, K., Kandasamy, S. and Narayanan, M. (2024) Influences of Wildfire on the Forest Ecosystem and Climate Change: A Comprehensive Study. Environmental Research, 240, Article 117537. [Google Scholar] [CrossRef] [PubMed]
[2] Lian, J., Pan, X. and Guo, J. (2023) An Improved Fire and Smoke Detection Method Based on YOLOv7. 2023 32nd International Conference on Computer Communications and Networks (ICCCN), Honolulu, 24-27 July 2023, 1-7. [Google Scholar] [CrossRef
[3] 易冠霖, 吴浩峻, 吴韵哲, 等. 基于语义分割的船舶机舱初期火灾识别算法[J]. 船海工程, 2023, 52(5): 109-114.
[4] 胡久松, 刘张驰, 余谦, 等. 融入GhostNet和CBAM的YOLOv8烟雾识别算法[J]. 电子测量与仪器学报, 2024, 38(8): 201-207.
[5] 李永福, 陈立斌, 惠君伟, 等. 基于EPSA-YOLOv5电力高空作业安全带佩戴检测[J]. 西安工程大学学报, 2024, 38(2): 18-25.
[6] 高均益, 张伟, 李泽麟. YOLO-BFEPS: 一种高效注意力增强的跨尺度YOLOv10火灾检测模型[J]. 计算机科学, 2025, 52(S1): 424-432.
[7] 曲英伟, 刘锐. 基于YOLOv5-MobileNetV3算法的目标检测[J]. 计算机系统应用, 2024, 33(6): 213-221.
[8] Su, L., Zhang, S. and Ding, W. (2023) An Improved Real-Time Detection Method for Flame and Smoke Identification Based on YOLOv5. 2023 6th International Conference on Intelligent Autonomous Systems (ICoIAS), Qinhuangdao, 22-24 September 2023, 59-64. [Google Scholar] [CrossRef
[9] 李敏学, 张晓宇, 程英杰, 等. FireLight-YOLO:面向森林火灾实时监测的轻量化模型[J]. 北京林业大学学报, 2026, 48(1): 12-25.
[10] Phan, D.T., Yap, K.H., Garg, K. and Han, B.S. (2023) Vision-Based Early Fire and Smoke Detection for Smart Factory Applications Using FFS-YOLO. 2023 IEEE 25th International Workshop on Multimedia Signal Processing (MMSP), Poitiers, 27-29 September 2023, 1-6. [Google Scholar] [CrossRef
[11] Redmon, J. and Farhadi, A. (2018) YOLOv3: An Incremental Improvement. arXiv: 1804.02767.
[12] Li, C., Li, L., Jiang, H., et al. (2022) YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv: 2209.02976.
[13] Wang, C., Bochkovskiy, A. and Liao, H.M. (2023) YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 7464-7475. [Google Scholar] [CrossRef
[14] Bochkovskiy, A., Wang, C.Y. and Liao, H.Y.M. (2020) YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv: 2004.10934.
[15] Wang, C.Y., Yeh, I.H. and Mark Liao, H. (2024) YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T. and Varol, G., Eds., Lecture Notes in Computer Science, Springer, 1-21. [Google Scholar] [CrossRef
[16] Khanam, R. and Hussain, M. (2024) YOLOv11: An Overview of the Key Architectural Enhancements. arXiv: 2410.17725.
[17] Wang, Z., Li, C., Xu, H., Zhu, X. and Li, H. (2025) Mamba YOLO: A Simple Baseline for Object Detection with State Space Model. Proceedings of the AAAI Conference on Artificial Intelligence, 39, 8205-8213. [Google Scholar] [CrossRef
[18] Li, Y., Hu, J., Wen, Y., Evangelidis, G., Salahi, K., Wang, Y., et al. (2023) Rethinking Vision Transformers for Mobilenet Size and Speed. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, 1-6 October 2023, 16889-16900. [Google Scholar] [CrossRef
[19] Jiao, J., Liu, Y., Liu, Y., Tian, Y., Wang, Y., Xie, L., et al. (2024) VMamba: Visual State Space Model. Advances in Neural Information Processing Systems, 37, 103031-103063. [Google Scholar] [CrossRef
[20] Lu, L.P., Xiong, Q., Xu, B. and Chu, D. (2024) MixDehazeNet: Mix Structure Block for Image Dehazing Network. 2024 International Joint Conference on Neural Networks (IJCNN), Yokohama, 30 June 2024-5 July 2024, 1-10. [Google Scholar] [CrossRef
[21] Dosovitskiy, A. (2020) An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv: 2010.11929.
[22] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., et al. (2021) Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 10012-10022. [Google Scholar] [CrossRef
[23] Nam, J., Syazwany, N.S., Kim, S.J. and Lee, S. (2024) Modality-Agnostic Domain Generalizable Medical Image Segmentation by Multi-Frequency in Multi-Scale Attention. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 16-22 June 2024, 11480-11491. [Google Scholar] [CrossRef
[24] Gu, A.R., Nam, J.H. and Lee, S.C. (2022) FBI-Net: Frequency-Based Image Forgery Localization via Multitask Learning with Self-attention. IEEE Access, 10, 62751-62762. [Google Scholar] [CrossRef
[25] Qin, Z., Zhang, P., Wu, F. and Li, X. (2021) FcaNet: Frequency Channel Attention Networks. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 783-792. [Google Scholar] [CrossRef
[26] Sang, M. and Hansen, J.H.L. (2022) Multi-Frequency Information Enhanced Channel Attention Module for Speaker Representation Learning. Proceedings of Interspeech 2022, Incheon, 18-22 September 2022, 321-325. [Google Scholar] [CrossRef
[27] Gu, A. and Dao, T. (2024) Mamba: Linear-Time Sequence Modeling with Selective State Space. arXiv: 2312.00752.
[28] Liu, Y., Shao, Z. and Hoffmann, N. (2021) Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions. arXiv: 2112.05561.
[29] Kelenyi, B., Domsa, V. and Tamas, L. (2024) Sam-Net: Self-Attention Based Feature Matching with Spatial Transformers and Knowledge Distillation. Expert Systems with Applications, 242, Article 122804. [Google Scholar] [CrossRef
[30] Hou, Q., Zhou, D. and Feng, J. (2021) Coordinate Attention for Efficient Mobile Network Design. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 20-25 June 2021, 13713-13722. [Google Scholar] [CrossRef
[31] Sandler, M., Howard, A., Zhu, M., Zhmoginov, A. and Chen, L. (2018) MobileNetV2: Inverted Residuals and Linear Bottlenecks. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, 18-23 June 2018, 4510-4520. [Google Scholar] [CrossRef