ChatGPT辅助大学生中医体质辨别的可行性研究
Feasibility of ChatGPT-Assisted TCM Constitution Classification for College Students
DOI: 10.12677/acm.2026.1682874, PDF,    科研立项经费支持
作者: 王 茜*:云南医药健康职业学院基础与药学院,云南 昆明;张 娟#:云南医药健康职业学院临床医学院,云南 昆明
关键词: 中医体质;体质辨别;ChatGPT;大学生;TCM Constitution; Constitution Classification; ChatGPT; College Students
摘要: 目的:基于云南医药健康职业学院500名大学生中医体质分布数据,评估ChatGPT-4.0模拟并复现《中医体质分类与判定表》判别逻辑的可行性与准确性。方法:以《中医体质分类与判定表》对500名大学生进行传统量表测评,计算9种体质转化分并判定主体质及兼夹体质,并将该量表测评结果作为参照标准;同时将量表症状描述转换为结构化提示词输入ChatGPT-4.0,分析其输出结果与传统量表转化分的相关性及判别效能。结果:500名大学生中偏颇体质者占71.20%,主倾向体质前三位为阳虚质(29.00%)、阴虚质(20.80%)、湿热质(18.80%),兼夹体质者占总人群50.00%;女性阳虚质、阴虚质、气虚质、血瘀质、特禀质、气郁质转化分均高于男性,平和质转化分低于男性(P < 0.05);ChatGPT-4.0总体判别准确率为81.20% (406/500),Kappa = 0.687 (P < 0.001);灵敏度居前三位为阳虚质(87.23%)、气虚质(85.42%)、平和质(83.33%),特异度居前三位为特禀质(95.89%)、血瘀质(94.12%)、痰湿质(92.94%);亚组分析中,无兼夹体质组判别准确率为90.40%,有兼夹体质组为72.00%,非明确偏颇组为84.39%,明确偏颇组为78.33%,两组间差异均有统计学意义(χ2 = 27.589, 5.927, P < 0.05)。结论:ChatGPT-4.0在复现《中医体质分类与判定表》的判别逻辑方面具有中等一致性(准确率81.20%),对单一偏颇体质识别效能较好,但在兼夹体质判别中准确率存在下降。本研究表明,大语言模型可作为高校中医体质量表化初步筛查的辅助工具,但鉴于量表本身作为参照标准存在固有局限性,目前尚不能替代专业中医师的综合判断。
Abstract: Objective: To evaluate the feasibility and accuracy of ChatGPT-4.0 in simulating and replicating the discriminative logic of the Classification and Determination of Constitution in Traditional Chinese Medicine Scale, based on the distribution data of TCM constitution types among 500 students at Yunnan Medical Health College. Methods: A total of 500 college students were assessed using the Classification and Determination of Constitution in Traditional Chinese Medicine Scale. Transformation scores for the nine constitution types were calculated to identify the dominant constitution and coexisting constitution types for each participant, and these scale-based classification results were adopted as the reference standard. Additionally, symptom descriptions from the scale were converted into structured prompts and fed into ChatGPT-4.0. The correlation between ChatGPT’s outputs and the transformation scores of the traditional scale, as well as its discriminative performance, were statistically analyzed. Results: Among the 500 college students, 71.20% presented with biased constitutions. The top three dominant types were yang-deficiency constitution (29.00%), yin-deficiency constitution (20.80%), and dampness-heat constitution (18.80%), with coexisting constitutions observed in 50.00% of all subjects. Females had significantly higher transformation scores for yang-deficiency constitution, yin-deficiency constitution, qi-deficiency constitution, blood-stasis constitution, special diathesis constitution, and qi-depression constitution, but a significantly lower score for balanced constitution compared with males (P < 0.05). The overall identification accuracy of ChatGPT-4.0 reached 81.20% (406/500), with a Kappa coefficient of 0.687 (P < 0.001). The top three types in terms of sensitivity were yang-deficiency constitution (87.23%), qi-deficiency constitution (85.42%), and balanced constitution (83.33%). The top three in specificity were special diathesis constitution (95.89%), blood-stasis constitution (94.12%), and phlegm-dampness constitution (92.94%). Subgroup analysis revealed that the identification accuracy was 90.40% in the group without coexisting constitutions, 72.00% in the group with coexisting constitutions, 84.39% in the non-obvious biased group, and 78.33% in the obvious biased group. The differences between the two subgroup pairs were statistically significant (χ2 = 27.589 and 5.927 respectively; both P < 0.05). Conclusion: ChatGPT-4.0 demonstrated moderate agreement (overall accuracy: 81.20%) in replicating the discriminative logic of the Classification and Determination of Constitution in Traditional Chinese Medicine Scale, with satisfactory performance in identifying single biased constitutions but decreased accuracy in cases involving coexisting constitutions. This study suggests that large language models may serve as an auxiliary tool for scale-based preliminary screening of TCM constitutions in university settings; however, given the inherent limitations of the scale itself as a reference standard, they are not yet capable of replacing the comprehensive judgment of professional TCM practitioners.
文章引用:王茜, 张娟. ChatGPT辅助大学生中医体质辨别的可行性研究[J]. 临床医学进展, 2026, 16(8): 1001-1008. https://doi.org/10.12677/acm.2026.1682874

参考文献

[1] 何汶霞, 王俊杰, 刘泽, 等. 郴州市某高校大学生中医体质兼夹规律调查[J]. 湘南学院学报(医学版), 2023, 25(4): 52-55.
[2] 边文彦, 徐学权. 广州大学生行为方式与中医体质相关性研究[J]. 光明中医, 2025, 40(5): 861-864.
[3] 中华中医药学会. 中医体质分类与判定[S]. 北京: 中国中医药出版社, 2009.
[4] 潘晨浩, 周明. ChatGPT在中医药领域中的应用前景[J]. 医学信息, 2023, 36(18): 15-20.
[5] 王琦, 朱燕波. 中国一般人群中医体质流行病学调查——基于全国9省市21948例流行病学调查数据[J]. 中华中医药杂志, 2009, 24(1): 7-12.
[6] 中医体质分类与判定(ZYYXH/T157-2009) [J]. 世界中西医结合杂志, 2009, 4(4): 303-304.
[7] https://d.wanfangdata.com.cn/periodical/CiBQZXJpb2RpY2FsQ0hJU29scjkyMDI2MDcyNzAzMTgxMhINZ3N6eTIwMjMxMTAyMRoIbTNma2s3cWQ%3D
[8] 李邱诺, 杨柯, 杨春红, 等. 1568例大学生中医体质调查与调护——以贵州中医药大学为例[J]. 贵州中医药大学学报, 2024, 46(2): 86-90.
[9] Hu, H., Yang, Q., Hu, J., Fan, Y., Zhang, Z., Yang, K., et al. (2026) ESRRA-ATG5-Mediated Mitophagy Enhances Arginine Metabolism to Alleviate Diabetic Kidney Disease. Autophagy, 22, 666-690.
https://doi.org/10.1080/15548627.2025.2601874
[10] Ma, G., Li, C., Ji, P., Chen, Y., Li, A., Hu, Q., et al. (2023) Association of Traditional Chinese Medicine Body Constitution and Cold Syndrome with Leukocyte Mitochondrial Functions: An Observational Study. Medicine, 102, e32694.
https://doi.org/10.1097/md.0000000000032694
[11] Liu, Y., Yuan, Y., Yan, K., Li, Y., Sacca, V., Hodges, S., et al. (2025) Evaluating the Role of Large Language Models in Traditional Chinese Medicine Diagnosis and Treatment Recommendations. npj Digital Medicine, 8, Article No. 466.
https://doi.org/10.1038/s41746-025-01845-2
[12] Hager, P., Jungmann, F., Holland, R., Bhagat, K., Hubrecht, I., Knauer, M., et al. (2024) Evaluation and Mitigation of the Limitations of Large Language Models in Clinical Decision-Making. Nature Medicine, 30, 2613-2622.
https://doi.org/10.1038/s41591-024-03097-1