基于强化学习与动态异构图的电商跨域商品序列推荐方法
A Reinforcement Learning and Dynamic Heterogeneous Graph-Based Method for Cross-Domain Product Sequential Recommendation in E-Commerce
DOI: 10.12677/ecl.2026.156641, PDF,   
作者: 王则林, 王 祥, 程 实:南通大学人工智能与计算机学院,江苏 南通;赵书娴*:南通大学人工智能与计算机学院,江苏 南通;江苏汇环环保科技有限公司,江苏 南通
关键词: 跨域序列推荐动态异构图离线强化学习迁移代价约束Cross-Domain Sequential Recommendation Dynamic Heterogeneous Graph Offline Reinforcement Learning Migration Cost Constraint
摘要: 本文面向电商领域跨域商品推荐场景,针对目标域商品交互数据稀疏、用户长期偏好难以有效建模与优化的问题,在动态异构图表征与多视图对齐学习的基础上,将推荐任务建模为带迁移代价约束的受限马尔可夫决策过程,提出离线强化学习方法DHGCRL。模型以源域局部图、目标域局部图与跨域全局异构图生成的三视图表征构造状态,并结合候选集约束与SlateQ列表价值近似实现Top-K商品列表决策。奖励函数由目标域商品反馈与跨域对齐相似度塑形项共同构成,以增强迁移表征对目标域商品偏好的刻画能力。同时,采用推荐列表表征与源域稳定兴趣表征之间的余弦距离定义迁移代价,通过联合回报价值与代价价值形成修正价值,从而指导策略学习并满足代价阈值约束。为提升离线训练的稳定性,进一步引入保守价值学习与行为一致性正则。实验结果表明,DHGCRL在HR和NDCG等指标上均优于多种序列推荐与离线强化学习基线方法。
Abstract: In the context of cross-domain product recommendation in e-commerce, this study addresses the sparsity of target-domain interaction data and the difficulty of modeling and optimizing users’ long-term preferences. Building on dynamic heterogeneous graph representations and multi-view alignment learning, the recommendation process is formulated as a constrained Markov decision process with transfer cost constraints, and an offline reinforcement learning method, DHGCRL, is proposed. The model constructs the state from three-view representations derived from the source-domain local graph, target-domain local graph, and cross-domain global heterogeneous graph, and performs Top-K product list decision-making by combining candidate set constraints with Slate Q-based list value approximation. The reward function integrates target-domain feedback with a cross-domain alignment similarity shaping term to strengthen target-domain preference modeling. Meanwhile, the transfer cost is defined as the cosine distance between the recommended list representation and the stable interest representation in the source domain. By jointly modeling return value and cost value, the method learns a revised value function that guides policy optimization under a predefined cost threshold. To improve the stability of offline training, conservative value learning and behavior-consistency regularization are further introduced. Experimental results demonstrate that DHGCRL outperforms a range of sequential recommendation and offline reinforcement learning baselines in terms of HR and NDCG.
文章引用:王则林, 王祥, 赵书娴, 程实. 基于强化学习与动态异构图的电商跨域商品序列推荐方法[J]. 电子商务评论, 2026, 15(6): 335-345. https://doi.org/10.12677/ecl.2026.156641

参考文献

[1] Chen, S., Xu, Z., Pan, W., et al. (2024) A Survey on Cross-Domain Sequential Recommendation. The 33rd International Joint Conference on Artificial Intelligence Survey Track, Jeju, 3-9 August 2024, 7989-7998.
[2] Lin, G., Gao, C., Zheng, Y., Chang, J., Niu, Y., Song, Y., et al. (2024) Mixed Attention Network for Cross-Domain Sequential Recommendation. Proceedings of the 17th ACM International Conference on Web Search and Data Mining, Merida, 4-8 March 2024, 405-413. [Google Scholar] [CrossRef
[3] Ni, R., Cai, W. and Jiang, Y. (2024) Contrastive Cross-Domain Sequential Recommendation via Emphasized Intention Features. Neural Networks, 179, Article ID: 106488. [Google Scholar] [CrossRef] [PubMed]
[4] Kang, W. and McAuley, J. (2018) Self-Attentive Sequential Recommendation. 2018 IEEE International Conference on Data Mining (ICDM), Singapore, 17-20 November 2018, 197-206. [Google Scholar] [CrossRef
[5] Sun, F., Liu, J., Wu, J., Pei, C., Lin, X., Ou, W., et al. (2019) BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, 3-7 November 2019, 1441-1450. [Google Scholar] [CrossRef
[6] Zhou, D., Cai, X. and Pan, W. (2025) Contrastive Text-Enhanced Transformer for Cross-Domain Sequential Recommendation. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Vol. 2, 4110-4119. [Google Scholar] [CrossRef
[7] Li, W., Lin, X., Pan, W. and Ming, Z. (2024) Dynamic Stage-Aware User Interest Learning for Heterogeneous Sequential Recommendation. 18th ACM Conference on Recommender Systems, Bari, 14-18 October 2024, 465-474. [Google Scholar] [CrossRef
[8] Liu, S., Cai, Q., He, Z., Sun, B., McAuley, J., Zheng, D., et al. (2023) Generative Flow Network for Listwise Recommendation. Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Long Beach, 6-10 August 2023, 1524-1534. [Google Scholar] [CrossRef
[9] Chen, X., Wang, S., McAuley, J., Jannach, D. and Yao, L. (2024) On the Opportunities and Challenges of Offline Reinforcement Learning for Recommender Systems. ACM Transactions on Information Systems, 42, 1-26. [Google Scholar] [CrossRef
[10] Wang, K., Zou, Z., Zhao, M., Deng, Q., Shang, Y., Liang, Y., et al. (2023) RL4RS: A Real-World Dataset for Reinforcement Learning Based Recommender System. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, Taipei, 23-27 July 2023, 2935-2944. [Google Scholar] [CrossRef
[11] Wachi, A., Shen, X. and Sui, Y. (2024) A Survey of Constraint Formulations in Safe Reinforcement Learning. Proceedings of the 33rd International Joint Conference on Artificial Intelligence, Jeju, 3-9 August 2024, 8262-8271.
[12] Chen, X., Yao, L., McAuley, J., Zhou, G. and Wang, X. (2023) Deep Reinforcement Learning in Recommender Systems: A Survey and New Perspectives. Knowledge-Based Systems, 264, Article ID: 110335. [Google Scholar] [CrossRef