人机协同视域下生成式人工智能辅助高中英语词汇选择题编写研究
A Study on Generative AI-Assisted Senior High School English Vocabulary Multiple-Choice Item Generation from the Perspective of Human-AI Collaboration
摘要: 在生成式人工智能技术快速发展的背景下,AI为语言测试开发提供了新的可能,但现有研究对教师如何运用AI进行人机协同命题的实践过程关注不足。本研究以DeepSeek为辅助工具,探索编写高中英语词汇选择题的可行路径与协作模式。研究发现,DeepSeek在试题格式规范、教材主题契合和语言表达方面具有较好的辅助价值,但AI生成试题仍存在考查内容覆盖不均衡、语境真实性与支持程度不足、正确答案唯一性和干扰项有效性不稳定等问题,需经人工审查与修改方可达到使用标准。研究建议,教师在使用AI辅助命题时应坚持教师主导、人机协同的原则,以提高英语词汇测试开发的效率与质量。
Abstract: Against the backdrop of rapid advances in generative artificial intelligence, AI has opened up new possibilities for language test development. However, existing research has paid insufficient attention to how teachers can engage in human-AI collaborative item generation in practice. This study employs DeepSeek as an assisting tool to explore feasible pathways and collaborative models to develop multiple-choice vocabulary test items for senior high school English. The findings reveal that DeepSeek demonstrates considerable value in terms of item formatting consistency, thematic alignment with textbook content, and naturalness of language expression. Nevertheless, AI-generated items still exhibit limitations, including uneven coverage of target vocabulary, inadequate contextual authenticity and support, and instability in answer uniqueness and distractor effectiveness, all of which necessitate human review and revision before the items can meet quality standards. The study recommends that teachers adhere to the principle of teacher leadership and human-AI collaboration when utilizing AI for item development, so as to enhance both the efficiency and quality of English vocabulary test development.
参考文献
|
[1]
|
中华人民共和国教育部. 普通高中英语课程标准(2017年版2025年修订) [S]. 北京: 人民教育出版社, 2025.
|
|
[2]
|
宋德龙. 高中英语词汇教学与测试策略的研究[J]. 中小学英语教学与研究, 2008, 31(8): 29-32.
|
|
[3]
|
Shin, D. and Lee, J.H. (2023) Can ChatGPT Make Reading Comprehension Testing Items on Par with Human Experts? Language Learning & Technology, 27, 27-40. https://doi.org/10.64152/10125/73530
|
|
[4]
|
Khademi, M. (2023) Can ChatGPT and Bard Generate Aligned Assessment Items? A Reliability Analysis against Human Performance. Journal of Applied Learning & Teaching, 6, 75-80.
|
|
[5]
|
Lin, Z. and Chen, H. (2024) Investigating the Capability of ChatGPT for Generating Multiple-Choice Reading Comprehension Items. System, 123, Article ID: 103344. https://doi.org/10.1016/j.system.2024.103344
|
|
[6]
|
Read, J. and Chapelle, C.A. (2001) A Framework for Second Language Vocabulary Assessment. Language Testing, 18, 1-32. https://doi.org/10.1177/026553220101800101
|
|
[7]
|
Heaton, J.B. (1975) Writing English Language Tests. Longman.
|
|
[8]
|
陈晓扣, 郑庆珠. 论英语词汇测试中的高低、一致原则[J]. 解放军外国语学院学报, 2000(6): 64-66.
|
|
[9]
|
乔辉. 情境化试题设计在高考英语中的应用探析[J]. 中小学外语教学(中学篇), 2024, 47(12): 55-61.
|
|
[10]
|
周艳琼. 人机协同命题与英语选择题阅读测试效度——基于IRT的多维度比较[J]. 外语教学与研究, 2026, 58(3): 430-443.
|
|
[11]
|
Messick, S. (1989) Validity. In: Linn, R.L., Ed., Educational Measurement (3rd Edition), Macmillan, 13-103.
|
|
[12]
|
Bachman, L.F. and Palmer, A.S. (1996) Language Testing in Practice: Designing and Developing Useful Language Tests. Oxford University Press.
|
|
[13]
|
中华人民共和国教育部, 国家语言文字工作委员会. GF 0018-2018中国英语能力等级量表[S]. 北京: 高等教育出版社, 2018.
|
|
[14]
|
Sweller, J. (1988) Cognitive Load during Problem Solving: Effects on Learning. Cognitive Science, 12, 257-285. https://doi.org/10.1207/s15516709cog1202_4
|