基于学生视角的大语言模型生成护理硕士跨学科连续照护模拟案例的评价研究
An Evaluation Study on the Generation of Interdisciplinary Continuous Care Simulation Cases for Master of Nursing by LLM from a Student Perspective
摘要: 目的:本研究旨在系统比较DeepSeek R1、GEMINI和豆包三种大语言模型(LLM)生成三幕式连续照护案例的质量差异。方法:采用标准化提示词驱动三款LLM各生成5例覆盖常见临床场景的案例,由23名护理硕士研究生基于验证过的评价指标体系进行盲法评价,涵盖案例结构、内容(含跨学科协作与循证思维)及应用效果维度。结果:DeepSeek R1、GEMINI和豆包三种模型在综合总分依次为(258.43 ± 67.68)、(245.96 ± 74.01)、(245.43 ± 74.56),无显著统计学差异(P > 0.788),且所有二级指标也均无统计学差异(P > 0.05);然而,跨学科协作深度仅达中等水平,在该维度的得分DeepSeek R1、GEMINI和豆包依次为(25.57 ± 7.04)、(24.87 ± 6.70)、(23.91 ± 7.67)。结论:本研究所测试的三种LLM在结构化提示下生成案例的整体质量趋于一致且基本满足教学需求,可作为高效的护理模拟案例辅助生成工具,但在跨学科互动真实性与伦理安全性上仍需人工审校。未来研究应聚焦提示词工程优化以增强情境深度,并通过真实教学场景试验验证其对学生临床推理能力的实际影响,从而构建“人机协同”的护理教育新范式。
Abstract: Objective: This study aims to systematically compare the quality differences in generating three-act continuous care case scenarios by three large language models (LLMs): DeepSeek R1, GEMINI, and DouBao. Methods: Standardized prompts were used to generate five cases each from the three LLMs, covering common clinical scenarios. Twenty-three master’s students in nursing conducted blinded evaluations using a validated assessment framework, evaluating case structure, content (including interdisciplinary collaboration and evidence-based thinking), and application effectiveness. Results: The overall composite scores of DeepSeek, GEMINI, and DouBao were 258.43 ± 67.68, 245.96 ± 74.01, and 245.43 ± 74.56, respectively, with no statistically significant differences (P > 0.788); all secondary indicators also showed no significant differences (P > 0.05). However, the depth of interdisciplinary collaboration was only moderate, with scores of 25.57 ± 7.04,24.87 ± 6.70, and 23.91 ± 7.67 for DeepSeek R1, GEMINI, and DouBao, respectively. Conclusion: Current mainstream LLMs produce cases of consistent overall quality under structured prompts, meeting basic teaching requirements and serving as efficient auxiliary tools for nursing simulation case generation. However, manual review remains necessary to ensure the authenticity of interdisciplinary interactions and ethical safety. Future research should focus on optimizing prompt engineering to enhance contextual depth and validate its practical impact on students’ clinical reasoning abilities through real-world teaching scenario experiments, thereby establishing a new paradigm of “human-machine collaboration” in nursing education.
文章引用:张桃桃, 程妍, 徐蔡洁, 商丽. 基于学生视角的大语言模型生成护理硕士跨学科连续照护模拟案例的评价研究[J]. 创新教育研究, 2026, 14(9): 807-814. https://doi.org/10.12677/ces.2026.149747

参考文献

[1] Yuan, Y. (2025) Impact of Multidisciplinary Continuity of Care on Postoperative Outcomes in Liver Cancer Surgical Patients. Journal of Multidisciplinary Healthcare, 18, 4749-4759.
https://doi.org/10.2147/jmdh.s527399
[2] 葛向煜, 贾守梅, 胡雁, 等. 健康信息方向护理硕士专业学位研究生培养的探索与实践[J]. 中华护理教育, 2024, 21(7): 828-834.
[3] Zhang, C., Xu, C., Wang, R., Han, X., Yang, G. and Liu, Y. (2024) The Learning Experiences and Career Development Expectations of Chinese Nursing Master’s Degree Students: A Qualitative Investigation. Nurse Education in Practice, 77, Article 103996.
https://doi.org/10.1016/j.nepr.2024.103996
[4] 郭欣怡, 李琨, 赵娟娟, 等. 护理硕士专业学位研究生高级健康评估的进阶式课程设计和多梯度仿真模拟教学实践[J]. 中国医学教育技术, 2024, 38(1): 81-86.
[5] Franco-Tantuico, M.A. and Btoush, R. (2025) Simulation Debriefing in Graduate Nursing Education: A Literature Review. Nursing Education Perspectives, 46, 278-283.
https://doi.org/10.1097/01.nep.0000000000001434
[6] Laurito, W., Davis, B., Grietzer, P., Gavenčiak, T., Böhm, A. and Kulveit, J. (2025) AI-AI Bias: Large Language Models Favor Communications Generated by Large Language Models. Proceedings of the National Academy of Sciences, 122, e241569712.
https://doi.org/10.1073/pnas.2415697122
[7] Thirunavukarasu, A.J., Ting, D.S.J., Elangovan, K., Gutierrez, L., Tan, T.F. and Ting, D.S.W. (2023) Large Language Models in Medicine. Nature Medicine, 29, 1930-1940.
https://doi.org/10.1038/s41591-023-02448-8
[8] Ruggiano, N., Sahoo, S., Brashear, A., Nwatu, U., Brunson, A., Noh, H., et al. (2026) Evaluating AI-Generated Geriatric Case Studies for Interprofessional Education: Systematic Analysis across 5 Platforms. JMIR Medical Education, 12, e83085.
https://doi.org/10.2196/83085
[9] 欧阳子瑶, 樊德净, 潘晓, 等. 护理硕士专业学位研究生教学临床案例库评价指标体系的构建[J]. 中华护理教育, 2023, 20(6): 665-671.
[10] 齐星伦, 姚一帆, 沈舒施, 等. 不同大语言模型肿瘤标志物报告解读性能评价[J]. 检验医学, 2025, 40(11): 1075-1081.
[11] 黄慧, 胡瑾瑜, 王晓宇, 等. 不同大型语言模型与不同水平医学专业人士回答眼科问题的对比研究[J]. 国际眼科杂志, 2024, 24(3): 458-462.
[12] 赵朋伟, 李干, 薛紫阳, 等. 以 ChatGPT-4o和DeepSeek-V3为代表的生成式人工智能在[12]肠造口患者教育支持中的对比研究[J]. 军事护理, 2025, 42(12): 71-74.
[13] Jahnke, M.N., O’Haver, J., Gupta, D., Hawryluk, E.B., Finelt, N., Kruse, L., et al. (2021) Care of Congenital Melanocytic Nevi in Newborns and Infants: Review and Management Recommendations. Pediatrics, 148, e2021051536.
https://doi.org/10.1542/peds.2021-051536
[14] Cao, W., Zhang, Q., Liu, J. and Liu, S. (2026) From Agents to Governance: Essential AI Skills for Clinicians in the Large Language Model Era. Journal of Medical Internet Research, 28, e86550.
https://doi.org/10.2196/86550
[15] Tong, K., McMahon, E., Reid-McDermott, B., Byrne, D. and Doherty, A.M. (2021) SafePsych: Improving Patient Safety by Delivering High-Impact Simulation Training on Rare and Complex Scenarios in Psychiatry. BMJ Open Quality, 10, e001533.
https://doi.org/10.1136/bmjoq-2021-001533
[16] Xue, H., Yuan, H., Li, G., Liu, J. and Zhang, X. (2021) Comparison of Team-Based Learning vs. Lecture-Based Teaching with Small Group Discussion in a Master’s Degree in Nursing Education Course. Nurse Education Today, 105, Article 105043.
https://doi.org/10.1016/j.nedt.2021.105043