基于多模型协同优化的空气质量指数预测研究——以福州市为例
Air Quality Index Prediction Based on Multi-Model Collaborative Optimization—A Case Study of Fuzhou
摘要: 空气质量指数(Air Quality Index, AQI)的精准预测对于空气污染防治和环境管理具有重要意义。本文以福州市为研究对象,收集2021年1月至2025年3月的AQI、六项主要空气污染物及气象数据,采用XGBoost和Pearson相关系数进行特征筛选,构建XGBoost、BP神经网络、BiLSTM、Prophet等单一预测模型,以及STL-XGBoost、BiLSTM-STL-XGBoost、BP-BiLSTM-Stacking等多模型融合预测模型,并利用支持向量机(SVM)开展空气质量等级分类研究。结果表明:BP神经网络在单一模型中预测性能最佳;BP-BiLSTM-Stacking融合模型综合性能最优,在AQI预测任务中取得最高的预测精度;优化后的SVM模型能够有效提升空气质量等级分类效果,尤其提高了轻度污染类别的识别准确率。研究结果表明,多模型融合方法能够充分发挥不同模型的互补优势,提高AQI预测的准确性和稳定性,可为空气质量预测及相关环境管理研究提供参考。
Abstract: Accurate prediction of the Air Quality Index (AQI) is of great importance for air pollution prevention, environmental management, and public health protection. Taking Fuzhou as the study area, this study collected daily AQI data, concentrations of six major air pollutants, and meteorological data from January 2021 to March 2025. XGBoost and Pearson correlation analysis were employed for feature selection. Four individual prediction models, including XGBoost, Back Propagation Neural Network (BPNN), Bidirectional Long Short-Term Memory (BiLSTM), and Prophet, as well as several multi-model fusion frameworks, including STL-XGBoost, BiLSTM-STL-XGBoost, and BP-BiLSTM-Stacking, were developed and evaluated. In addition, a Support Vector Machine (SVM) was applied to classify air quality levels. The results indicate that the BPNN achieved the best performance among the individual models, while the BP-BiLSTM-Stacking model demonstrated the highest overall prediction accuracy for AQI forecasting. Moreover, the optimized SVM significantly improved the classification performance of air quality levels, particularly in identifying the minority class of lightly polluted days. The findings suggest that multi-model fusion can effectively exploit the complementary strengths of different models, thereby improving the accuracy and robustness of AQI prediction and providing a useful reference for air quality forecasting and environmental management.
参考文献
|
[1]
|
李子熠, 张天宇, 李鸿强. 基于XGBoost和LSTM组合模型的PM2.5浓度预测[J]. 河北建筑工程学院学报, 2023, 41(4): 219-223.
|
|
[2]
|
王子航, 王子睿, 柳卓行, 杨娟. 基于LSTM多步预测模型的空气质量预测与预警[J]. 应用数学进展, 2023, 12(12): 5057-5071.
|
|
[3]
|
Smola, A.J. and Schölkopf, B. (2004) A Tutorial on Support Vector Regression. Statistics and Computing, 14, 199-222. https://doi.org/10.1023/b:stco.0000035301.49549.88
|
|
[4]
|
Chen, T. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, 13-17 August 2016, 785-794. https://doi.org/10.1145/2939672.2939785
|
|
[5]
|
Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. Neural Computation, 9, 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735
|
|
[6]
|
Wolpert, D.H. (1992) Stacked Generalization. Neural Networks, 5, 241-259. https://doi.org/10.1016/s0893-6080(05)80023-1
|
|
[7]
|
纪峰, 陈荔, 李占利. 基于STL文件的模型及应用[J]. 长安大学学报(自然科学版), 2006, 26(1): 104-107.
|
|
[8]
|
常恬君, 过仲阳, 徐丽丽. 基于Prophet-随机森林优化模型的空气质量指数规模预测[J]. 环境污染与防治, 2019, 41(7): 758-761, 766.
|
|
[9]
|
李英英, 纪昌杰. 基于信息熵加权去噪的半监督SVM分类器[J]. 电脑知识与技术, 2013, 9(25): 5705-5707, 5715.
|