|
[1]
|
本刊编辑部. 第57次《中国互联网络发展状况统计报告》在京发布[J]. 中国教工, 2026(2): 66.
|
|
[2]
|
Brown, T., Mann, B., Ryder, N., et al. (2020) Language Models Are Few-Shot Learners. Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, 6-12 December 2020, 1877-1901.
|
|
[3]
|
Touvron, H., Lavril, T., Izacard, G., et al. (2023) LLaMA: Open and Efficient Foundation Language Models. https://arxiv.org/abs/2302.13971
|
|
[4]
|
Bai, J., Bai, S., Chu, Y., et al. (2023) Qwen Technical Report. https://arxiv.org/abs/2309.16609
|
|
[5]
|
Papineni, K., Roukos, S., Ward, T. and Zhu, W. (2002) BLEU: A Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, 6-12 July 2002, 311-318. https://doi.org/10.3115/1073083.1073135
|
|
[6]
|
Lin, C.Y. (2004) Rouge: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out, Barcelona, 25 July 2004, 74-81.
|
|
[7]
|
Vaswani, A., et al. (2017) Attention Is All You Need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, 4-9 December 2017, 6000-6010.
|
|
[8]
|
Zeng, A., Liu, X., Du, Z., et al. (2022) GLM-130B: An Open Bilingual Pre-Trained Model. https://arxiv.org/abs/2210.02414
|
|
[9]
|
Liu, A., Feng, B., Xue, B., et al. (2024) DeepSeek-V3 Technical Report. https://arxiv.org/abs/2412.19437
|
|
[10]
|
王思宇, 邱江涛, 洪川洋, 等. 基于知识图谱的在线商品问答研究[J]. 中文信息学报, 2020, 34(11): 104-112.
|
|
[11]
|
Zhang, S., Dinan, E., Urbanek, J., Szlam, A., Kiela, D. and Weston, J. (2018) Personalizing Dialogue Agents: I Have a Dog, Do You Have Pets Too? Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1, 2204-2213. https://doi.org/10.18653/v1/p18-1205
|
|
[12]
|
冯源, 钱松荣, 陆宇亮. 基于RAG本地电商知识库的DeepSeek电商模型构建与优化研究[J]. 电子商务评论, 2025, 14(5): 1346-1359.
|
|
[13]
|
Deriu, J., Rodrigo, A., Otegi, A., Echegoyen, G., Rosset, S., Agirre, E., et al. (2021) Survey on Evaluation Methods for Dialogue Systems. Artificial Intelligence Review, 54, 755-810. https://doi.org/10.1007/s10462-020-09866-x
|
|
[14]
|
Wang, C., Liu, X., Yue, Y., et al. (2023) Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity. https://arxiv.org/abs/2310.07521
|
|
[15]
|
Banerjee, S. and Lavie, A. (2005) METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Ann Arbor, June 2005, 65-72.
|
|
[16]
|
Zhang, T., Kishore, V., Wu, F., et al. (2019) BERTScore: Evaluating Text Generation with BERT. https://arxiv.org/abs/1904.09675
|
|
[17]
|
Sellam, T., Das, D. and Parikh, A. (2020) BLEURT: Learning Robust Metrics for Text Generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5-10 July 2020, 7881-7892. https://doi.org/10.18653/v1/2020.acl-main.704
|
|
[18]
|
Yuan, W., Neubig, G. and Liu, P. (2021) Bartscore: Evaluating Generated Text as Text Generation. Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, 6-14 December 2021, 27263-27277.
|
|
[19]
|
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W., Koh, P., et al. (2023) FActScore: Fine-Grained Atomic Evaluation of Factual Precision in Long Form Text Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6-10 December 2023, 12076-12100. https://doi.org/10.18653/v1/2023.emnlp-main.741
|
|
[20]
|
Fu, J., Ng, S., Jiang, Z. and Liu, P. (2024) GPTScore: Evaluate as You Desire. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, 6556-6576. https://doi.org/10.18653/v1/2024.naacl-long.365
|
|
[21]
|
Lin, Y.T. and Chen, Y.N. (2023) LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023), Toronto, 14 July 2023, 47-58.
|
|
[22]
|
Singhal, K., Tu, T., Gottweis, J., et al. (2023) Towards Expert-Level Medical Question Answering with Large Language Models. https://arxiv.org/abs/2305.09617
|
|
[23]
|
许建峰, 刘程远, 况琨, 何浩, 孙常龙, 李宝善, 等. 法律大模型评估指标和测评方法[J]. 中国人工智能学会通讯, 2024, 2(14): 10-22.
|