生成式人工智能与教师英语写作任务评改一致度研究
Consistency between GAI’s and Teachers’ Rating and Corrective Feedback in English Writing Task
摘要: 本研究旨在探讨生成式人工智能与教师在评改英语写作任务时的一致程度。共15名大学生参加了为期三个月的雅思写作任务2教学实验,由两位教师和两种生成式人工智能(KIMI与ChatGPT 4.0)分别对120份作文进行独立评分并提供文字版评改建议。通过计算生成式人工智能与教师评分的一致性,发现两种生成式人工智能评分的内部一致性高于两位教师的内部一致性;人工智能与教师间的一致程度一般。从成绩分布情况看,生成式人工智能成绩分布呈现比较集中的趋势,且成绩评定偏高于学生实际英语写作能力。通过编码分析教师与生成式人工智能对学生英语作文的评改建议,发现两种生成式人工智能在对120份作文的评改建议中稳定且一致地遵守了雅思作文2的评改维度;而两位教师则有时会有主观倾向,评改建议不能覆盖雅思作文2的四个评分维度。研究还发现,教师与生成式人工智能在评改建议提出方式方面也有差别。研究结论表明目前教师尚不能完全依赖人工智能评改作文,但可与人工智能协作开展写作任务评改,提高效率的同时,有效避免因教师人为因素产生的不一致性。
Abstract: This study aims to explore the consistency between generative artificial intelligence (GAI) and teachers in assessing English writing tasks. A total of 15 university students participated in a three-month teaching experiment focused on IELTS Writing Task 2. Two teachers and two generative AI systems (KIMI and ChatGPT 4.0) independently scored 120 essays and provided written feedback. By calculating the consistency between the AI and teacher scores, it was found that the internal consistency of the two AI systems was higher than that of the two teachers, while the consistency between the AI and teachers was moderate. In terms of score distribution, the AI systems exhibited a relatively concentrated trend, with scores generally higher than the students’ actual English writing proficiency. Through coding and analysis of the feedback provided by teachers and AI on the students’ essays, it was observed that the two AI systems consistently adhered to the assessment criteria of IELTS Writing Task 2 across all 120 essays. In contrast, the teachers occasionally displayed subjective tendencies and did not fully cover all four scoring dimensions of the task. Additionally, differences were noted in the manner in which teachers and AI delivered their feedback. The study concludes that, at present, teachers cannot fully rely on AI for essay assessment. However, collaboration between teachers and AI in evaluating writing tasks can enhance efficiency while effectively mitigating inconsistencies and subjectivity arising from human factors.
参考文献
|
[1]
|
Swain, M. (2005) The Output Hypothesis: Theory and Research. In: Hinkel, E., Ed., Handbook of Research in Second Language Teaching and Learning, Lawrence Erlbaum Associates, 471-483.
|
|
[2]
|
张雪梅. 大学英语写作教学现状之调查[J]. 外语界, 2006(5): 28-32.
|
|
[3]
|
Ding, L. and Zou, D. (2024) Automated Writing Evaluation Systems: A Systematic Review of Grammarly, Pigai, and Criterion with a Perspective on Future Directions in the Age of Generative Artificial Intelligence. Education and Information Technologies, 29, 14151-14203. https://doi.org/10.1007/s10639-023-12402-3
|
|
[4]
|
Li, X. and Zhang, L. (2019) A Critical Review of Pigai: Challenges and Opportunities in Automated Writing Assessment. Computer Assisted Language Learning, 32, 589-605.
|
|
[5]
|
陈茉, 吕明臣. ChatGPT环境下的大学英语写作教学[J]. 当代外语研究, 2024(1): 161-168.
|
|
[6]
|
任伟, 刘远博, 解月. 英语写作教学中ChatGPT与教师反馈的对比研究[J]. 外语教学理论与实践, 2024(4): 30-38, 60.
|
|
[7]
|
谭智. 应用Rasch模型分析英语写作评分行为[J]. 外语教学理论与实践, 2008(1): 26-31.
|
|
[8]
|
杨丽萍, 辛涛. 人工智能辅助能力测量: 写作自动化评分研究的核心问题[J]. 现代远程教育研究, 2021, 33(4): 51-62.
|
|
[9]
|
Zhang, Y. and Hyland, K. (208) Feedback in Second Language Writing: Contexts and Issues. Cambridge University Press. https://doi.org/10.1017/9781108635547
|
|
[10]
|
Warschauer, M. and Ware, P. (2006) Automated Writing Evaluation: Defining the Classroom Research Agenda. Language Teaching Research, 10, 157-180. https://doi.org/10.1191/1362168806lr190oa
|