机器学习预测中老年高血压风险的研究
Prediction of Hypertension Risk in Middle-Aged and Elderly Populations Using Machine Learning
摘要: 针对当前中老年高血压风险预测研究中模型泛化性不足、缺乏多方法对比与外部验证支撑等局限,本文基于CHARLS中老年队列数据与NHANES独立外部验证数据,构建“双重特征筛选–多模型对比–跨人群验证”的机器学习预测框架,为中老年高血压风险分层筛查提供科学依据。该方法整合人口学、体格测量、血压指标、生活方式与慢性病变量,采用单因素Logistic回归与Lasso回归联合筛选,剔除多重共线性指标后确定13项核心变量;构建Logistic回归、随机森林与XGBoost三种模型,以ROC曲线、AUC等指标评估性能,结合5折交叉验证与SHAP值解析验证模型稳健性与可解释性。结果表明,XGBoost模型综合性能最优,内部AUC达0.942,外部AUC为0.830,内外AUC差值仅0.112,无明显过拟合;年龄、BMI、收缩压、舒张压为核心危险因素。本研究构建的XGBoost模型兼具最优区分效能与跨人群泛化稳定性,可为中老年高血压早期筛查与个体化防控提供科学参考。此外,基于非血压指标的补充分析模型同样展现出良好的早期预警效能,与主模型形成分层互补,为中老年高血压的社区初筛提供了无创化技术方案。
Abstract: This paper presents a machine learning prediction framework of “dual feature selection-multi-model comparison-cross-population validation” to address insufficient generalizability and lack of external validation in current hypertension risk prediction models for middle-aged and elderly populations. Based on CHARLS cohort data and independent NHANES validation data, we integrated multi-dimensional variables, such as demographic, anthropometric, blood pressure, lifestyle, and chronic disease variables, and performed joint feature selection using univariate Logistic regression and Lasso regression, identifying 13 core variables. Three models—Logistic Regression, Random Forest, and XGBoost—were constructed. Performance was evaluated using ROC curves and AUC, and robustness and interpretability were verified using 5-fold cross-validation and SHAP value analysis. Results showed that the XGBoost model exhibited the best overall performance, with an internal AUC of 0.942 and an external AUC of 0.830, a difference of only 0.112, indicating no significant overfitting. Age, BMI, systolic blood pressure, and diastolic blood pressure were identified as the core risk factors. The XGBoost model constructed in this study combines optimal discriminative power with cross-population generalization stability, providing a scientific reference for early screening and individualized prevention of hypertension in middle-aged and elderly populations. Furthermore, the supplementary analysis model based on non-blood pressure indicators also demonstrated good early warning efficacy, forming a hierarchical complement to the main model and providing a non-invasive technical solution for community-based initial screening of hypertension in middle-aged and elderly populations.
文章引用:冉娟, 陈家旺, 殷爽, 李凌霄. 机器学习预测中老年高血压风险的研究[J]. 应用数学进展, 2026, 15(7): 263-278. https://doi.org/10.12677/aam.2026.157321

参考文献

[1] 张冬燕, 李燕. 世界卫生组织《全球高血压报告》(2023年)概要及解读[J]. 诊断学理论与实践, 2024, 23(3): 297-304.
[2] 刘嘉慧, 刘靖. 《中国高血压防治指南(2024年修订版)》亮点、要点解读[J]. 中国医学前沿杂志(电子版), 2025, 17(1): 1-6.
[3] 钱晰彦, 汤一帆, 李婷茹, 等. 高血压患者发生心血管疾病相关危险因素的Meta分析[J]. 中国循证心血管医学杂志, 2024, 16(12): 1416-1423.
[4] 中国高血压联盟《高血压患者高质量血压管理中国专家建议》委员会. 高血压患者高质量血压管理中国专家建议[J]. 中华高血压杂志(中英文), 2024, 32(2): 104-111.
[5] 曲帅, 张盈, 陈雪, 等. 高血压患病风险预测模型的研究进展[J]. 心脏杂志, 2025, 37(2): 216-220+235.
[6] 魏诗意, 张珍, 王叶婷, 等. 基于机器学习构建高血压患病风险预测模型的系统评价[J]. 牡丹江医学院学报, 2024, 45(5): 55-61.
[7] 沈赛拉, 钟锋, 梁兴, 等. 基于随机森林和梯度提升决策树的高血压分析预测[J]. 计算机时代, 2023(5): 15-19.
[8] 张添玉, 邢春国, 马雨杨, 等. 基于深度学习的高血压风险预测模型的构建[J]. 医药高职教育与现代护理, 2026, 9(2): 172-178.
[9] Xu, T., Liu, J., Liu, J., et al. (2021) Risk Factors for Hypertension in Chinese Adults: A Systematic Review and Meta-Analysis. Journal of Hypertension, 39, 1110-1119.
[10] Zhao, D., Liu, J., Wang, M., Zhang, X. and Zhou, M. (2019) Epidemiology of Cardiovascular Disease in China: Current Features and Implications. Nature Reviews Cardiology, 16, 203-212. [Google Scholar] [CrossRef] [PubMed]
[11] Lundberg, S.M. and Lee, S.I. (2017) A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, 4-9 December 2017, 4765-4774.
[12] 中国高血压防治指南修订委员会. 中国高血压防治指南(2024年修订版) [J]. 中国心血管杂志, 2024, 29(1): 1-56.
[13] Williams, B., Mancia, G., Spiering, W., Agabiti Rosei, E., Azizi, M., Burnier, M., et al. (2018) 2018 ESC/ESH Guidelines for the Management of Arterial Hypertension. European Heart Journal, 39, 3021-3104. [Google Scholar] [CrossRef] [PubMed]
[14] Riley, R.D., Ensor, J., Snell, K.I.E., Debray, T.P.A., Altman, D.G., Moons, K.G.M., et al. (2016) External Validation of Clinical Prediction Models Using Big Datasets from E-Health Records or IPD Meta-Analysis: Opportunities and Challenges. BMJ, 353, i3140. [Google Scholar] [CrossRef] [PubMed]
[15] Youden, W.J. (1950) Index for Rating Diagnostic Tests. Cancer, 3, 32-35. [Google Scholar] [CrossRef] [PubMed]
[16] Topol, E.J. (2019) High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nature Medicine, 25, 44-56. [Google Scholar] [CrossRef] [PubMed]
[17] Wiens, J., Saria, S., Sendak, M., Ghassemi, M., Liu, V.X., Doshi-Velez, F., et al. (2019) Do No Harm: A Roadmap for Responsible Machine Learning for Health Care. Nature Medicine, 25, 1337-1340. [Google Scholar] [CrossRef] [PubMed]
[18] Niiranen, T.J., McCabe, E.L., Larson, M.G., Henglin, M., Lakdawala, N.K., Vasan, R.S., et al. (2017) Risk for Hypertension Crosses Generations in the Community: A Multi-Generational Cohort Study. European Heart Journal, 38, 2300-2308. [Google Scholar] [CrossRef] [PubMed]
[19] Chen, T. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, 13-17 August 2016, 785-794. [Google Scholar] [CrossRef