CEENet:结合对比学习与视觉嵌入的情绪识别方法
CEENet: Emotion Recognition with Contrastive Learning and Visual Embeddings
摘要: 针对当前上下文情绪识别研究中抽象情绪语义学习不足、上下文线索融合方式单一以及面部表情特征分离质量不佳等困难,文章提出了一种结合对比学习与视觉嵌入增强的情绪识别网络CEENet (Context Embedding Emotion Network)。该网络通过并行提取背景、身体、面部以及基于视觉–语言模型生成的视觉嵌入共四路特征作为情绪识别的上下文,进一步加深了对高级情绪语义的挖掘。为实现多来源上下文特征的高效融合,文章设计了多流特征注意力模块,利用自适应注意力机制动态调整各上下文信息在情绪识别中的权重,强化情绪相关线索并抑制噪声。此外,针对面部表情特征的分离,文章提出了一种多尺寸特征增强的面部表情嵌入网络,并引入一种创新的三层级数据增强三元表情对比学习方案,显著提升了表情嵌入的判别力。实验结果表明,CEENet在Emotic和CAER-S等上下文情绪识别基准测试上分别达到了34.15%的平均精确率均值和88.25%的分类准确率,同时本研究中的面部表情嵌入网络也在FEC基准测试中取得了86.1%的准确率,均优于同类型主流方法。最后,消融实验进一步验证了主干特征提取结构、特征融合结构及多层级数据增强策略在提升模型性能方面的有效性。
Abstract: To address the challenges in current context-aware emotion recognition research, including inadequate learning of abstract emotional semantics, limited strategies for fusing contextual cues, and unsatisfactory disentanglement of facial expression features, this paper proposes CEENet (Context Embedding Emotion Network), an emotion recognition network that integrates contrastive learning with visual embedding enhancement. CEENet extracts four parallel streams of features as contextual cues for emotion recognition, namely background, body, face, and visual embeddings generated by a vision-language model, thereby enhancing the modeling of high-level emotional semantics. To achieve efficient fusion of multi-source contextual features, a Multi-Stream Feature Attention module is designed to dynamically adjust the weights of different contextual information through an adaptive attention mechanism, thereby strengthening emotion-relevant cues while suppressing noise. In addition, to address the disentanglement of facial expression features, this paper proposes a facial expression embedding network with multi-scale feature enhancement and introduces an innovative three-level data-augmentation-based triplet expression contrastive learning scheme. The proposed method significantly improves the discriminability of expression embeddings. Experimental results show that CEENet achieves a mean average precision of 34.15% and a classification accuracy of 88.25% on context-aware emotion recognition benchmarks such as Emotic and CAER-S, respectively. Meanwhile, the facial expression embedding network developed in this study attains an accuracy of 86.1% on the FEC benchmark, outperforming existing mainstream methods. Finally, ablation studies further verify the effectiveness of the backbone network design, the fusion network design, and the multi-level contrastive learning strategy in improving emotion recognition performance.
文章引用:夏旭辉, 田春岐. CEENet:结合对比学习与视觉嵌入的情绪识别方法[J]. 计算机科学与应用, 2026, 16(7): 154-168. https://doi.org/10.12677/csa.2026.167249

参考文献

[1] Ekman, P. and Friesen, W.V. (1971) Constants across Cultures in the Face and Emotion. Journal of Personality and Social Psychology, 17, 124-129. [Google Scholar] [CrossRef] [PubMed]
[2] Matsumoto, D. (1992) More Evidence for the Universality of a Contempt Expression. Motivation and Emotion, 16, 363-368. [Google Scholar] [CrossRef
[3] Bettadapura, V. (2012) Face Expression Recognition and Analysis: The State of the Art. arXiv:1203.6722.
[4] Saxena, A., Khanna, A. and Gupta, D. (2020) Emotion Recognition and Detection Methods: A Comprehensive Survey. Journal of Artificial Intelligence and Systems, 2, 53-79. [Google Scholar] [CrossRef
[5] Kosti, R., Alvarez, J.M., Recasens, A. and Lapedriza, A. (2017) EMOTIC: Emotions in Context Dataset. 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, 21-26 July 2017, 61-69. [Google Scholar] [CrossRef
[6] Vemulapalli, R. and Agarwala, A. (2019) A Compact Embedding for Facial Expression Similarity. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, 15-20 June 2019, 5683-5692. [Google Scholar] [CrossRef
[7] Zhang, W., Ji, X., Chen, K., Ding, Y. and Fan, C. (2021) Learning a Facial Expression Embedding Disentangled from Identity. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 20-25 June 2021, 6759-6768. [Google Scholar] [CrossRef
[8] Goodfellow, I.J., Erhan, D., Carrier, P.L., et al. (2013) Challenges in Representation Learning: A Report on Three Machine Learning Contests. Neural Information Processing: 20th International Conference, Daegu, 3-7 November 2013, 117-124.
[9] Ahmed, F., Bari, A.S.M.H. and Gavrilova, M.L. (2019) Emotion Recognition from Body Movement. IEEE Access, 8, 11761-11781. [Google Scholar] [CrossRef
[10] Cai, Y., Li, X. and Li, J. (2023) Emotion Recognition Using Different Sensors, Emotion Models, Methods and Datasets: A Comprehensive Review. Sensors, 23, Article 2455. [Google Scholar] [CrossRef] [PubMed]
[11] Mehrabian, A. and Russell, J.A. (1974) An Approach to Environmental Psychology. The MIT Press.
[12] Ninh, Q.B., Nguyen, H.C., Huynh, T., Tran, M. and Le, T. (2023) Multi-Branch Network for Imagery Emotion Prediction. Proceedings of the 12th International Symposium on Information and Communication Technology, Ho Chi Minh, 7-8 December 2023, 371-378. [Google Scholar] [CrossRef
[13] Mittal, T., Guhan, P., Bhattacharya, U., Chandra, R., Bera, A. and Manocha, D. (2020) Emoticon: Context-Aware Multimodal Emotion Recognition Using Frege’s Principle. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 14234-14243. [Google Scholar] [CrossRef
[14] Mittal, T., Bera, A. and Manocha, D. (2021) Multimodal and Context-Aware Emotion Perception Model with Multiplicative Fusion. IEEE MultiMedia, 28, 67-75. [Google Scholar] [CrossRef
[15] Radford, A., Kim, J.W., Hallacy, C., et al. (2021) Learning Transferable Visual Models from Natural Language Supervision. International Conference on Machine Learning, Virtual, 18-24 July 2021, 8748-8763.
[16] Li, B., Weinberger, K.Q., Belongie, S., et al. (2022) Language-Driven Semantic Segmentation. arXiv:2201.03546.
[17] Gu, X., Lin, T.Y., Kuo, W., et al. (2021) Open-Vocabulary Object Detection via Vision and Language Knowledge Distillation. arXiv:2104.13921.
[18] Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T. and Dubnov, S. (2023) Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, 4-10 June 2023, 1-5. [Google Scholar] [CrossRef
[19] Wang, C., Chai, M., He, M., Chen, D. and Liao, J. (2022) CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance Fields. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, 18-24 June 2022, 3835-3844. [Google Scholar] [CrossRef
[20] Li, J., Li, D., Xiong, C., et al. (2022) BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation. International Conference on Machine Learning, Baltimore, 17-23 July 2022, 12888-12900.
[21] Li, J., Li, D., Savarese, S., et al. (2023) BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. International Conference on Machine Learning, Honolulu, 23-29 July 2023 19730-19742.
[22] Zhang, J., Qu, X., Zhu, T. and Cheng, Y. (2025) CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, 4-9 November 2025, 5406-5419. [Google Scholar] [CrossRef
[23] Bai, S., Chen, K., Liu, X., et al. (2025) Qwen2. 5-vl Technical Report. arXiv:2502.13923.
[24] Zhang, S., Pan, Y. and Wang, J.Z. (2023) Learning Emotion Representations from Verbal and Nonverbal Communication. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 18993-19004. [Google Scholar] [CrossRef
[25] He, K., Zhang, X., Ren, S. and Sun, J. (2016) Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, 27-30 June 2016, 770-778. [Google Scholar] [CrossRef
[26] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., et al. (2021) Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 10012-10022. [Google Scholar] [CrossRef
[27] Parkhi, O.M., Vedaldi, A. and Zisserman, A. (2015) Deep Face Recognition. In: Xie, X., Jones, M.W. and Tam, G.K.L., Eds., Procedings of the British Machine Vision Conference 2015, BMVA Press, 41.1-41.12. [Google Scholar] [CrossRef
[28] Lee, J., Kim, S., Kim, S., Park, J. and Sohn, K. (2019) Context-Aware Emotion Recognition Networks. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, 27 October 2019-2 November 2019, 10143-10152. [Google Scholar] [CrossRef
[29] Li, W., Dong, X. and Wang, Y. (2021) Human Emotion Recognition with Relational Region-Level Analysis. IEEE Transactions on Affective Computing, 14, 650-663. [Google Scholar] [CrossRef
[30] Zhang, M., Liang, Y. and Ma, H. (2019) Context-Aware Affective Graph Reasoning for Emotion Recognition. 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, 8-12 July 2019, 151-156. [Google Scholar] [CrossRef
[31] Qing, L., Wen, H., Chen, H., Jin, R., Cheng, Y. and Peng, Y. (2024) DVC-Net: A New Dual-View Context-Aware Network for Emotion Recognition in the Wild. Neural Computing and Applications, 36, 653-665. [Google Scholar] [CrossRef
[32] Li, X., Peng, X. and Ding, C. (2021) Sequential Interactive Biased Network for Context-Aware Emotion Recognition. 2021 IEEE International Joint Conference on Biometrics (IJCB), Shenzhen, 4-7 August 2021, 1-6. [Google Scholar] [CrossRef
[33] Yuan, Y., Lu, F., Cheng, X. and Liu, Y. (2022) Context Based Vision Emotion Recognition in the Wild. 2022 IEEE 17th Conference on Industrial Electronics and Applications (ICIEA), Chengdu, 16-19 December 2022, 479-484. [Google Scholar] [CrossRef
[34] Gao, Q., Zeng, H., Li, G. and Tong, T. (2021) Graph Reasoning-Based Emotion Recognition Network. IEEE Access, 9, 6488-6497. [Google Scholar] [CrossRef
[35] Yang, D., Yang, K., Li, M., Wang, S., Wang, S. and Zhang, L. (2024) Robust Emotion Recognition in Context Debiasing. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 16-22 June 2024, 12447-12457. [Google Scholar] [CrossRef
[36] Nguyen, H., Tran, N., Nguyen, M., Ta, P. and D. Nguyen, H. (2026) Attention Mechanisms for Context-Aware Emotion Recognition. CCF Transactions on Pervasive Computing and Interaction, 8, 253-269. [Google Scholar] [CrossRef