深度伪造人脸图像检测方法研究进展:从强监督到无监督学习
Advances in Deepfake Facial Image Detection Methods: From Fully Supervised to Unsupervised Learning
摘要: 随着生成式人工智能快速发展,深度伪造人脸图像的真实感不断提升,对金融安全、司法取证和社会信任构成威胁,深度伪造人脸图像检测因此成为数字内容安全领域的重要研究方向。本文首先分析自编码器、生成对抗网络、语义可控生成和扩散模型等生成技术演进对检测线索变化的影响;其次按照训练过程中对标注信息的依赖程度,将现有检测方法归纳为强监督检测、弱监督与半监督检测、无监督检测三类,并总结其技术特点与局限。综述表明,相关研究正由显性伪影识别转向隐性特征建模,由闭集二分类转向跨域泛化与开放场景适应。未来需重点提升隐性特征挖掘、低标注学习、跨域泛化、复杂扰动鲁棒性和可解释评测能力,并关注视觉基础模型、多模态一致性检测与主动防御技术的发展。
Abstract: With the rapid development of generative artificial intelligence, deepfake facial images have become increasingly realistic, posing threats to financial security, judicial forensics, and social trust. Deepfake facial image detection has therefore become an important research direction in digital content security. This paper first analyzes how the evolution of generative techniques, including autoencoders, generative adversarial networks, semantically controllable generation, and diffusion models, has changed detectable forgery traces. Then, according to the dependence on annotation information during training, existing detection methods are categorized into three groups: fully supervised detection, weakly supervised and semi-supervised detection, and unsupervised detection, with their technical characteristics and limitations summarized. This review shows that related research is shifting from explicit artifact recognition to implicit feature modeling, and from closed-set binary classification to cross-domain generalization and open-scenario adaptation. Future research should focus on improving implicit feature mining, low-annotation learning, cross-domain generalization, robustness against complex perturbations, and interpretable evaluation, while also paying attention to the development of vision foundation models, multimodal consistency detection, and active defense techniques.
文章引用:魏雅霓, 王丹琳. 深度伪造人脸图像检测方法研究进展:从强监督到无监督学习[J]. 计算机科学与应用, 2026, 16(7): 37-46. https://doi.org/10.12677/csa.2026.167239

参考文献

[1] Cole, S. (2018) We Are Truly Fucked: Everyone Is Making AI-Generated Fake Porn Now. Motherboard/Vice.
https://www.vice.com/en/article/reddit-fake-porn-app-daisy-ridley/
[2] Perov, I., Gao, D., Chervoniy, N., et al. (2020) DeepFaceLab: Integrated, Flexible and Extensible Face-Swapping Framework. arXiv:2005.05535.
[3] Shahzad, H.F., Rustam, F., Flores, E.S., Luís Vidal Mazón, J., de la Torre Diez, I. and Ashraf, I. (2022) A Review of Image Processing Techniques for Deepfakes. Sensors, 22, Article 4556. [Google Scholar] [CrossRef] [PubMed]
[4] Faceswap Team (2017) Faceswap. GitHub.
https://github.com/deepfakes/faceswap
[5] Li, L., Bao, J., Yang, H., Chen, D. and Wen, F. (2020) Advancing High Fidelity Identity Swapping for Forgery Detection. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 5074-5083. [Google Scholar] [CrossRef
[6] Nirkin, Y., Keller, Y. and Hassner, T. (2019) FSGAN: Subject Agnostic Face Swapping and Reenactment. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, 27 October 2019-2 November 2019, 7184-7193. [Google Scholar] [CrossRef
[7] Chen, R., Chen, X., Ni, B. and Ge, Y. (2020) SimSwap: An Efficient Framework for High Fidelity Face Swapping. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, 12-16 October 2020, 2003-2011. [Google Scholar] [CrossRef
[8] Shaoanlu (2026) Faceswap-GAN.
https://github.com/shaoanlu/faceswap-GAN
[9] Tang, F.X. (2019) Chinese ‘Deepfake’ App Censured over Privacy Concerns. Sixth Tone.
https://www.sixthtone.com/news/1004513
[10] FaceApp Technology Limited (2026) FaceApp: Perfect Face Editor.
https://www.faceapp.com/
[11] Deng, Y., Yang, J., Chen, D., Wen, F. and Tong, X. (2020) Disentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 5154-5163. [Google Scholar] [CrossRef
[12] Tewari, A., Elgharib, M., Bharaj, G., Bernard, F., Seidel, H., Perez, P., et al. (2020) StyleRig: Rigging StyleGAN for 3D Control over Portrait Images. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 6142-6151. [Google Scholar] [CrossRef
[13] Ren, Y., Li, G., Chen, Y., Li, T.H. and Liu, S. (2021) PiRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 13759-13768. [Google Scholar] [CrossRef
[14] Snap Inc. (2026) Face Swap Component. Snap for Developers.
https://developers.snap.com/lens-studio/features/ar-tracking/face/face-swap
[15] https://reface.ai/
[16] Kim, K., Kim, Y., Cho, S., Seo, J., Nam, J., Lee, K., et al. (2025) Diffface: Diffusion-Based Face Swapping with Facial Guidance. Pattern Recognition, 163, Article 111451. [Google Scholar] [CrossRef
[17] Zhao, W., Rao, Y., Shi, W., Liu, Z., Zhou, J. and Lu, J. (2023) Diffswap: High-Fidelity and Controllable Face Swapping via 3d-Aware Masked Diffusion. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 8568-8577. [Google Scholar] [CrossRef
[18] Ding, Z., Zhang, X., Xia, Z., Jebe, L., Tu, Z. and Zhang, X. (2023) DiffusionRig: Learning Personalized Priors for Facial Appearance Editing. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-24 June 2023, 12736-12746. [Google Scholar] [CrossRef
[19] Han, Y., Zhu, J., He, K., Chen, X., Ge, Y., Li, W., et al. (2024) Face-Adapter for Pre-Trained Diffusion Models with fine-Grained ID and attribute Control. In: Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T. and Varol, G., Eds., Lecture Notes in Computer Science, Springer, 20-36. [Google Scholar] [CrossRef
[20] FaceFusion Team (2026) FaceFusion. GitHub.
https://github.com/facefusion/facefusion
[21] Afchar, D., Nozick, V., Yamagishi, J. and Echizen, I. (2018) MesoNet: A Compact Facial Video Forgery Detection Network. 2018 IEEE International Workshop on Information Forensics and Security (WIFS), Hong Kong, China, 11-13 December 2018, 1-7. [Google Scholar] [CrossRef
[22] Chollet, F. (2017) Xception: Deep Learning with Depthwise Separable Convolutions. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, 21-26 July 2017, 1251-1258. [Google Scholar] [CrossRef
[23] Qian, Y., Yin, G., Sheng, L., Chen, Z. and Shao, J. (2020) Thinking in Frequency: Face Forgery Detection by Mining Frequency-Aware Clues. In: Vedaldi, A., Bischof, H., Brox, T. and Frahm, JM., Eds., Lecture Notes in Computer Science, Springer International Publishing, 86-103. [Google Scholar] [CrossRef
[24] Xu, B., Liu, J., Liang, J., Lu, W. and Zhang, Y. (2021) Deepfake Videos Detection Based on Texture Features. Computers, Materials & Continua, 68, 1375-1388. [Google Scholar] [CrossRef
[25] Nataraj, L., Mohammed, T.M., Chandrasekaran, S., et al. (2019) Detecting GAN Generated Fake Images Using Co-Occurrence Matrices. arXiv:1903.06836.
[26] Barni, M., Kallas, K., Nowroozi, E. and Tondi, B. (2020) CNN Detection of GAN-Generated Face Images Based on Cross-Band Co-Occurrences Analysis. 2020 IEEE International Workshop on Information Forensics and Security (WIFS), New York, 6-11 December 2020, 1-6. [Google Scholar] [CrossRef
[27] Cozzolino, D. and Verdoliva, L. (2020) Noiseprint: A CNN-Based Camera Model Fingerprint. IEEE Transactions on Information Forensics and Security, 15, 144-159. [Google Scholar] [CrossRef
[28] Li, L., Bao, J., Zhang, T., Yang, H., Chen, D., Wen, F., et al. (2020) Face X-Ray for More General Face Forgery Detection. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 5001-5010. [Google Scholar] [CrossRef
[29] Zhao, H., Wei, T., Zhou, W., Zhang, W., Chen, D. and Yu, N. (2021) Multi-Attentional Deepfake Detection. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 20-25 June 2021, 2185-2194. [Google Scholar] [CrossRef
[30] Yang, Z., Tao, R., Zhang, C., et al. (2025) Leveraging Unlabeled Data from Unknown Sources via Dual-Path Guidance for Deepfake Face Detection. arXiv:2508.09022.
[31] Chen, L., Zhang, Y., Song, Y., Liu, L. and Wang, J. (2022) Self-Supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detection. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, 18-24 June 2022, 18710-18719. [Google Scholar] [CrossRef
[32] Ruff, L., Vandermeulen, R., Goernitz, N., et al. (2018) Deep One-Class Classification. International Conference on Machine Learning, 80, 4393-4402.
[33] Khalid, H. and Woo, S.S. (2020) OC-Fakedect: Classifying Deepfakes Using One-Class Variational Autoencoder. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, 14-19 June 2020, 656-657. [Google Scholar] [CrossRef
[34] Chen, B. and Tan, S. (2021) Featuretransfer: Unsupervised Domain Adaptation for Cross-Domain Deepfake Detection. Security and Communication Networks, 2021, Article ID: 9942754. [Google Scholar] [CrossRef
[35] Wang, Q., Wang, X., Liu, Z., Bai, N., Zhao, M. and Pang, S. (2026) Unsupervised Domain Adaptation-Based Cross-Type Deepfake Image Detection. IEEE Transactions on Image Processing, 35, 4411-4424. [Google Scholar] [CrossRef
[36] Radford, A., Kim, J.W., Hallacy, C., et al. (2021) Learning Transferable Visual Models from Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research, 139, 8748-8763.
[37] Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., et al. (2021) Emerging Properties in Self-Supervised Vision Transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, 10-17 October 2021, 9650-9660. [Google Scholar] [CrossRef
[38] Coalition for Content Provenance and Authenticity (2026) Content Credentials: C2PA Technical Specification, Version 2.2.
https://spec.c2pa.org/specifications/specifications/2.2/specs/C2PA_Specification.html