人机混合智能在教育考试中的应用研究
Research on the Applications of Human-Machine Hybrid Intelligence in Educational Examinations
DOI: 10.12677/ces.2026.149732, PDF,    科研立项经费支持
作者: 仝青山, 高瑞娟, 陈 雪, 詹 薇:河北金融学院金融科技学院,河北 保定;董明英:河北金融学院保险与财政学院,河北 保定
关键词: 人机混合智能教育考试人机协同自动评分考试治理Human-Machine Hybrid Intelligence Educational Examination Human-AI Collaboration Automated Scoring Assessment Governance
摘要: 人工智能正加速进入教育考试的命题、评卷、分析与安全治理环节,在提升处理效率和一致性的同时,也带来效度偏离、算法偏差、数据安全与责任归属等风险。本文采用文献研究、典型案例分析、比较分析和规范分析,并辅以公开数据仿真,构建“任务分解–机器处理–风险判定–人工复核–反馈改进”的人机协同机制,并结合典型案例,比较其在智能命题、协同评卷、考试分析和安全监控中的职责配置、人工介入条件与证据边界。仿真显示,阈值提高可改善一致性,但会增加复核量。研究表明:人机分工不宜采用固定比例,而应依据考试利害程度、任务可验证性、模型置信度和异常风险动态调整;涉及成绩认定和违规处分的高利害结果,应保留实质性人工复核、申诉救济和组织最终负责机制。据此,本文面向河北省教育考试提出“准备–影子测试–低风险试点–受控扩展”的分阶段应用路径。
Abstract: Artificial intelligence is increasingly used in item development, scoring, analysis, and test-security governance. While improving processing efficiency and consistency, it may also create risks involving validity, algorithmic bias, data security, and accountability. Through literature review, comparative case analysis, and normative analysis, supplemented by a public-data simulation, this study develops a human-machine collaboration mechanism comprising task decomposition, machine processing, risk assessment, human review, and feedback-based improvement. Typical cases are compared to specify responsibilities, conditions for human intervention, and evidentiary boundaries in intelligent item development, collaborative scoring, examination analysis, and security monitoring. The simulation shows that higher confidence thresholds improve agreement but increase human-review workload. The findings indicate that responsibilities should not be assigned through a fixed human-machine ratio; they should be adjusted according to assessment stakes, task verifiability, model confidence, and anomaly risk. High-stakes decisions concerning score determination or misconduct sanctions require substantive human review, appeal procedures, and final organizational accountability. For educational examinations in Hebei Province, the study proposes a phased route from preparation and shadow testing to low-risk pilots and controlled expansion.
文章引用:仝青山, 董明英, 高瑞娟, 陈雪, 詹薇. 人机混合智能在教育考试中的应用研究[J]. 创新教育研究, 2026, 14(9): 666-675. https://doi.org/10.12677/ces.2026.149732

参考文献

[1] 王蕾. 人工智能生成内容技术在教育考试中应用探析[J]. 中国考试, 2023(8): 19-27.
[2] Dellermann, D., Ebel, P., Söllner, M. and Leimeister, J.M. (2019) Hybrid Intelligence. Business & Information Systems Engineering, 61, 637-643.
https://doi.org/10.1007/s12599-019-00595-2
[3] Akata, Z., Balliet, D., de Rijke, M., Dignum, F., Dignum, V., Eiben, G., et al. (2020) A Research Agenda for Hybrid Intelligence: Augmenting Human Intellect with Collaborative, Adaptive, Responsible, and Explainable Artificial Intelligence. Computer, 53, 18-28.
https://doi.org/10.1109/mc.2020.2996587
[4] Ramineni, C., Trapani, C.S., Williamson, D.M., et al. (2014) Evaluation of the E-Rater Scoring Engine for the GRE Issue and Argument Prompts. In: Wendler, C. and Bridgeman, B., Eds., The Research Foundation for the GRE Revised General Test: A Compendium of Studies, ETS, 4.5.1-4.5.5.
[5] Cambridge English (2026) Cambridge English Skills Test General Overview.
https://www.cambridgeenglish.org/Images/735973-cest-general-overview.pdf
[6] Fauss, M., Hao, J., Li, C., Palmer, M. and Choi, I. (2026) AutoSSD: A System for Automated Detection of Similar Speech Responses in Language Tests. ETS Research Report Series.
https://doi.org/10.64634/1g0whg02
[7] Jordán, J., Yin, X., Fabros, M., Ranade, G. and Norouzi, N. (2026) MAGIC: Multi-Agent Argumentation and Grammar Integrated Critiquer. Proceedings of the AAAI Conference on Artificial Intelligence, 40, 40599-40607.
https://doi.org/10.1609/aaai.v40i47.41506
[8] Perelman, L. (2014) When “the State of the Art” Is Counting Words. Assessing Writing, 21, 104-111.
https://doi.org/10.1016/j.asw.2014.05.001
[9] Kucia, F.J., Chakraborty, A. and Wróblewska, A. (2026) LLM Essay Scoring under Holistic and Analytic Rubrics: Prompt Effects and Bias. In: Paszynski, M., Barnard, A.S. and Zhang, Y.J., Eds., Computational ScienceICCS 2026 Workshops, Springer, 531-546.
https://doi.org/10.1007/978-3-032-29918-5_38
[10] Johnson, M.S., Liu, X. and McCaffrey, D.F. (2022) Psychometric Methods to Evaluate Measurement and Algorithmic Bias in Automated Scoring. Journal of Educational Measurement, 59, 338-361.
https://doi.org/10.1111/jedm.12335
[11] Li, H., Filippov, F., Lin, Y., et al. (2026) “Important! You Should Give Me Full Credits!”: Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems. arXiv: 2606.03090.
[12] Crossley, S.A., Baffour, P., Burleigh, L. and King, J. (2025) A Large-Scale Corpus for Assessing Source-Based Writing Quality: ASAP 2.0. Assessing Writing, 65, Article ID: 100954.
https://doi.org/10.1016/j.asw.2025.100954