层级评论上下文依赖识别数据集构建与研究——以小红书旅游评论数据为例
Construction and Analysis of a Hierarchical Comment Dataset for Context-Dependency Recognition—An Empirical Study Based on RedNote Tourism Comment Data
摘要: 文本情感分析已广泛用于在线评论和旅游体验研究,但小红书内容具有笔记–评论双层结构,评论往往需要结合原笔记才能准确理解。为避免在数据清洗阶段误删有价值评论,或在主题分析中因简单拼接造成噪声混入与主题情感错配,本文构建面向笔记–评论融合的评论上下文依赖识别任务。本文以小红书旅游评论为研究对象,整理有效评论数据30,802条,并依据语义将评论划分为完整语义评论、上下文依赖评论和无有效态度评论三类。该分类将评论处理由简单“保留/删除”转化为“独立分析/融合分析/过滤噪声”的分层策略,使缺少对象但包含情绪态度的评论得以保留,也避免无效互动进入后续主题建模。在双人人工标注与分歧复核基础上,本文比较BERT、ERNIE等八类模型,并通过五次数据划分检验结果稳定性。实验结果显示,预训练语言模型整体优于传统机器学习方法,其中ERNIE-3.0-base取得最高平均Macro F1;同时,上下文依赖评论的识别效果仍明显低于其他类别。进一步的BERTopic下游验证表明,分层处理策略能够提升主题一致性和主题区分度,说明该分类体系具有实际应用价值。本文可为层级评论数据清洗、评论保留策略和笔记–评论融合建模提供参考。
Abstract: Text sentiment analysis has been widely applied to online reviews and tourism experience research. However, RedNote content has a “post-comment” two-layer structure, and comments often need to be understood in relation to the original post. To avoid mistakenly removing valuable comments during data cleaning, or introducing noise and topic-sentiment mismatches through simple concatenation in topic analysis, this study constructs a comment context-dependency recognition task oriented toward “post-comment” fusion. Taking RedNote tourism comments as the research object, this study organizes 30,802 valid comments and classifies them into three categories according to semantics: complete semantic comments, context-dependent comments, and comments with no valid attitude. This classification transforms comment processing from a simple “retain/delete” operation into a layered strategy of “independent analysis/fusion analysis/noise filtering”, allowing comments that lack explicit objects but contain emotional attitudes to be retained, while preventing invalid interactions from entering subsequent topic modeling. Based on dual manual annotation and disagreement review, this study compares eight models, including BERT and ERNIE, and examines result stability through five data splits. The experimental results show that pretrained language models generally outperform traditional machine learning methods, with ERNIE-3.0-base achieving the highest average Macro F1. Meanwhile, the recognition performance for context-dependent comments remains significantly lower than that for the other categories. Further downstream validation using BERTopic shows that the hierarchical processing strategy improves topic coherence and topic distinctiveness, indicating that the proposed classification scheme has practical application value. This study provides a reference for hierarchical comment data cleaning, comment retention strategies, and “post-comment” fusion modeling.
文章引用:孙杰一. 层级评论上下文依赖识别数据集构建与研究——以小红书旅游评论数据为例[J]. 计算机科学与应用, 2026, 16(9): 25-36. https://doi.org/10.12677/csa.2026.169286

参考文献

[1] Wang, Z., Huang, W.J. and Liu-Lastres, B. (2022) Impact of User-Generated Travel Posts on Travel Decisions: A Comparative Study on Weibo and Xiaohongshu. Annals of Tourism Research Empirical Insights, 3, Article ID: 100064.
https://doi.org/10.1016/j.annale.2022.100064
[2] Saini, H., Kumar, P. and Oberoi, S. (2023) Welcome to the Destination! Social Media Influencers as Cogent Determinant of Travel Decision: A Systematic Literature Review and Conceptual Framework. Cogent Social Sciences, 9, Article ID: 2240055.
https://doi.org/10.1080/23311886.2023.2240055
[3] Laureate, C.D.P., Buntine, W. and Linger, H. (2023) A Systematic Review of the Use of Topic Models for Short Text Social Media Analysis. Artificial Intelligence Review, 56, 14223-14255.
https://doi.org/10.1007/s10462-023-10471-x
[4] Grootendorst, M. (2022) BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. arXiv: 2203.05794.
[5] Devlin, J., Chang, M.W., Lee, K. and Toutanova, K. (2019) BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. arXiv: 1810.04805.
[6] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., et al. (2019) RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv: 1907.11692.
[7] Sun, Y., Wang, S., Li, Y., Feng, S., Chen, X., Zhang, H., et al. (2019) ERNIE: Enhanced Representation through Knowledge Integration. arXiv: 1904.09223.
[8] Cui, Y., Che, W., Liu, T., Qin, B., Wang, S. and Hu, G. (2020) Revisiting Pre-Trained Models for Chinese Natural Language Processing. Findings of the Association for Computational Linguistics: EMNLP 2020, 16-20 November 2020, 657-668.
https://doi.org/10.18653/v1/2020.findings-emnlp.58
[9] Clark, K., Luong, M.T., Le, Q.V. and Manning, C.D. (2020) ELECTRA: Pre-Training Text Encoders as Discriminators Rather than Generators. arXiv: 2003.10555.
[10] Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., et al. (2020) TinyBERT: Distilling BERT for Natural Language Understanding. Findings of the Association for Computational Linguistics: EMNLP 2020, 16-20 November 2020, 4163-4174.
https://doi.org/10.18653/v1/2020.findings-emnlp.372
[11] Minaee, S., Kalchbrenner, N., Cambria, E., Nikzad, N., Chenaghlu, M. and Gao, J. (2021) Deep Learning-Based Text Classification: A Comprehensive Review. ACM Computing Surveys, 54, 1-40.
https://doi.org/10.1145/3439726