基于文本相似度的搜索推荐点击预测模型
Improvement of the Recommended Click Prediction Model Based on Text Similarity
DOI: 10.12677/CSA.2019.93069, PDF,    科研立项经费支持
作者: 詹 彬, 吴晓鸰*, 凌 捷:广东工业大学计算机学院,广东 广州
关键词: 搜索推荐搜索点击预测词向量点击模型关键字语义Search Recommendation Search Click Prediction Term Vectors Click Model Keyword Semantics
摘要: 为了进一步提高用户在搜索引擎中的推荐内容点击预测准确率,本文采用了一种包含搜索内容相似度特征的方法。该方法的结构是由多个决策树构成的预测模型,使用了层次化softmax (Hierarchical softmax)将结果转换为二分类结果。为了理解用户搜索文本的语义,使用用户输入与推荐内容标题、高频相关内容以及推荐内容标签的文本相似度来增加点击预测模型的准确率。采用jieba分词将文本中的词汇切分出来,使用word2cev对所有词汇进行训练,构建词向量模型。最后使用LightGBM进行预测模型的构建。从205万条用户搜索记录中取出5万条作为验证集,剩下的作为训练集。实验结果表明,添加相似度特征之后模型的点击预测准确率得到了提升。
Abstract: In order to further improve the accuracy of the recommended content click prediction in search engine, a method based on the similarity feature of search content is proposed. The structure of the method is composed of multiple decision tree models. The Hierarchical softmax is used to convert the result to binary classification results. In order to understand the semantics of the user’s search text, the user input text similarity with the recommended content title, high-frequency related content, and recommended content tags is used to increase the accuracy of the click prediction model. The segmentation of words in the text is performed using the jieba segmentation, and word2cev is used to train all the words and construct a word vector model. Finally, Light GBM is used to build the prediction model. Then, 50,000 of the 2.05 million users’ search records are taken as the verification set and the rest as the training set. Experimental results show that the accuracy of the model is improved after adding similarity features.
文章引用:詹彬, 吴晓鸰, 凌捷. 基于文本相似度的搜索推荐点击预测模型[J]. 计算机科学与应用, 2019, 9(3): 613-621. https://doi.org/10.12677/CSA.2019.93069

参考文献

[1] 李晓明, 闫宏飞, 王继民. 搜索引擎: 原理、技术与系统[J]. 2012.
[2] Yang, M.C., Lee, D.G., Park, S.Y., et al. (2015) Knowledge-Based Question Answering Using the Semantic Embedding Space. Expert Systems with Applications, 42, 9086-9104. [Google Scholar] [CrossRef
[3] Joachims, T. (2002) Optimizing Search Engines Using Clickthrough Data. ACM Conference on Knowledge Discovery & Data Mining, Edmonton, 23-26 July 2002, 1-21.
[4] Xing, Q., Liu, Y., Nie, J.Y., et al. (2013) Incorporating User Preferences into Click Models.
[5] 汉语信息处理词汇01部分: 基本术语(GB12200.1-90)6 [S]. 北京: 中国标准出版社, 1991.
[6] Forney, G.D. (1993) The Viterbi Algorithm. Proceedings of the IEEE, 61, 268-278.
[7] Hinton, G.E. (1989) Learning Distributed Representations of Concepts. 8th Conference of the Cognitive Science Society, Ann Arbor, 1989, 1-11.
[8] Manning, C.D. (1999) Foundations of Statistical Natural Language Processing. MIT Press, Cambridge.
[9] Bengio, Y., Schwenk, H., Senécal, J., et al. (2003) Neural Probabilistic Language Models. Journal of Machine Learning Research, 3, 1137-1155.
[10] Lee, K., Park, C., Kim, N., et al. (2018) Accelerating Recurrent Neural Network Language Model Based Online Speech Recognition System.
[11] Deng, H., Lei, Z. and Wang, L. (2017) Global Context-Dependent Recurrent Neural Network Lan-guage Model with Sparse Feature Learning. Neural Computing & Applications, No. 6, 1-13.
[12] Shao, T., Chen, H. and Chen, W. (2018) Query Auto-Completion Based on Word2vec Semantic Similarity. Journal of Physics Conference Series, 1004, Article ID: 012018. [Google Scholar] [CrossRef
[13] 周练. Word2vec的工作原理及应用探究[J]. 图书情报导刊, 2015(2): 145-148.
[14] Kearns, M.J. and Valiant, L.G. (1993) Cryptographic Limitations on Learning Boolean Formulae and Finite Automata. Springer-Verlag, Berlin. [Google Scholar] [CrossRef
[15] Valiant, L. (2015) Probably Approximately Correct: Nature’s Algorithms for Learning and Prospering in a Complex World. Common Knowledge, 21, 340. [Google Scholar] [CrossRef
[16] Guolin, K., Qing, M. and Thomas, F. (2017) LightGBM: A Highly Efficient Gradient Boosting Decision Tree. 31st Conference on Neural Information Processing Systems, Long Beach, 2017, 1-11.
[17] Shi, H. (2007) Best-First Decision Tree Learning. The University of Waikato, Hillcrest.