|
[1]
|
王蕾. 人工智能生成内容技术在教育考试中应用探析[J]. 中国考试, 2023(8): 19-27.
|
|
[2]
|
Dellermann, D., Ebel, P., Söllner, M. and Leimeister, J.M. (2019) Hybrid Intelligence. Business & Information Systems Engineering, 61, 637-643. https://doi.org/10.1007/s12599-019-00595-2
|
|
[3]
|
Akata, Z., Balliet, D., de Rijke, M., Dignum, F., Dignum, V., Eiben, G., et al. (2020) A Research Agenda for Hybrid Intelligence: Augmenting Human Intellect with Collaborative, Adaptive, Responsible, and Explainable Artificial Intelligence. Computer, 53, 18-28. https://doi.org/10.1109/mc.2020.2996587
|
|
[4]
|
Ramineni, C., Trapani, C.S., Williamson, D.M., et al. (2014) Evaluation of the E-Rater Scoring Engine for the GRE Issue and Argument Prompts. In: Wendler, C. and Bridgeman, B., Eds., The Research Foundation for the GRE Revised General Test: A Compendium of Studies, ETS, 4.5.1-4.5.5.
|
|
[5]
|
Cambridge English (2026) Cambridge English Skills Test General Overview. https://www.cambridgeenglish.org/Images/735973-cest-general-overview.pdf
|
|
[6]
|
Fauss, M., Hao, J., Li, C., Palmer, M. and Choi, I. (2026) AutoSSD: A System for Automated Detection of Similar Speech Responses in Language Tests. ETS Research Report Series. https://doi.org/10.64634/1g0whg02
|
|
[7]
|
Jordán, J., Yin, X., Fabros, M., Ranade, G. and Norouzi, N. (2026) MAGIC: Multi-Agent Argumentation and Grammar Integrated Critiquer. Proceedings of the AAAI Conference on Artificial Intelligence, 40, 40599-40607. https://doi.org/10.1609/aaai.v40i47.41506
|
|
[8]
|
Perelman, L. (2014) When “the State of the Art” Is Counting Words. Assessing Writing, 21, 104-111. https://doi.org/10.1016/j.asw.2014.05.001
|
|
[9]
|
Kucia, F.J., Chakraborty, A. and Wróblewska, A. (2026) LLM Essay Scoring under Holistic and Analytic Rubrics: Prompt Effects and Bias. In: Paszynski, M., Barnard, A.S. and Zhang, Y.J., Eds., Computational Science—ICCS 2026 Workshops, Springer, 531-546. https://doi.org/10.1007/978-3-032-29918-5_38
|
|
[10]
|
Johnson, M.S., Liu, X. and McCaffrey, D.F. (2022) Psychometric Methods to Evaluate Measurement and Algorithmic Bias in Automated Scoring. Journal of Educational Measurement, 59, 338-361. https://doi.org/10.1111/jedm.12335
|
|
[11]
|
Li, H., Filippov, F., Lin, Y., et al. (2026) “Important! You Should Give Me Full Credits!”: Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems. arXiv: 2606.03090.
|
|
[12]
|
Crossley, S.A., Baffour, P., Burleigh, L. and King, J. (2025) A Large-Scale Corpus for Assessing Source-Based Writing Quality: ASAP 2.0. Assessing Writing, 65, Article ID: 100954. https://doi.org/10.1016/j.asw.2025.100954
|