基于原型对比学习的可解释放射学报告生成方法
Interpretable Radiology Report Generation via Prototype-Guided Contrastive Learning
DOI: 10.12677/hjbm.2026.165093, PDF,    科研立项经费支持
作者: 张颜秀祺, 李慧敏*:云南民族大学数学与计算机科学学院,云南 昆明
关键词: 放射学报告生成;原型学习;对比学习;跨模态对齐;Radiology Report Generation; Prototype Learning; Contrastive Learning; Cross-Modal Alignment
摘要: 自动放射学报告生成的目标是根据胸部X射线图像自动写出临床描述,以减轻放射科医生的报告撰写负担。已有方法虽取得一定进展,但生成结果缺乏可解释性,同时图像特征与文本语义之间没有显式对齐,限制了报告质量的提升。为此,我们提出ProtoR2Gen框架,引入一组可学习的疾病原型作为语义锚点,利用对比学习将图像特征向对应的原型聚集,从而在视觉编码和文本生成之间建立明确的对应关系。在推理时,原型的激活分布可转化为图像上的热力图,直观展示模型关注的解剖区域。在IU X-Ray数据集上,ProtoR2Gen的CIDEr指标达到0.656,显著优于对比方法。消融实验和可视化结果也证明了原型机制对语义对齐和可解释性的提升作用。
Abstract: Automatic radiology report generation aims to automatically produce clinical descriptions from chest X‑ray images to reduce the documentation burden on radiologists. Existing methods have made some progress, but the generated results lack interpretability, and visual features are not explicitly aligned with textual semantics, which limits further improvement in report quality. To address this, we propose ProtoR2Gen, a framework that introduces a set of learnable disease prototypes as semantic anchors. Through contrastive learning, image features are pulled toward their corresponding prototypes, establishing a clear correspondence between visual encoding and text generation. During inference, the activation distribution of prototypes can be converted into heatmaps over the image, intuitively showing which anatomical regions the model focuses on. On the IU X‑Ray dataset, ProtoR2Gen achieves a CIDEr score of 0.656, significantly outperforming comparison methods. Ablation studies and visualization results also demonstrate the effectiveness of the prototype mechanism in improving semantic alignment and interpretability.
文章引用:张颜秀祺, 李慧敏. 基于原型对比学习的可解释放射学报告生成方法[J]. 生物医学, 2026, 16(5): 914-924. https://doi.org/10.12677/hjbm.2026.165093

参考文献

[1] Shin, H.C., Roberts, K., Lu, L., Demner-Fushman, D., Yao, J. and Summers, R.M. (2016) Learning to Read Chest X-Rays: Recurrent Neural Cascade Model for Automated Image Annotation. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, 27-30 June 2016, 2497-2506.
https://doi.org/10.1109/cvpr.2016.274
[2] Jing, B., Xie, P. and Xing, E. (2018) On the Automatic Generation of Medical Imaging Reports. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1, 2577-2586.
https://doi.org/10.18653/v1/p18-1240
[3] Li, Y., Liang, X., Hu, Z., et al. (2018) Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation. Advances in Neural Information Processing Systems, Montréal, 3-8 December 2018, 1537-1547.
[4] Chen, Z., Song, Y., Chang, T. and Wan, X. (2020) Generating Radiology Reports via Memory-Driven Transformer. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 16-20 November 2020, 1439-1449.
https://doi.org/10.18653/v1/2020.emnlp-main.112
[5] Chen, Z., Shen, Y., Song, Y. and Wan, X. (2021) Cross-Modal Memory Networks for Radiology Report Generation. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Volume 1, 5904-5914.
https://doi.org/10.18653/v1/2021.acl-long.459
[6] Cornia, M., Stefanini, M., Baraldi, L. and Cucchiara, R. (2020) Meshed-Memory Transformer for Image Captioning. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 10575-10584.
https://doi.org/10.1109/cvpr42600.2020.01059
[7] Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., et al. (2024) VMamba: Visual State Space Model. Advances in Neural Information Processing Systems 37, Vancouver, 9-15 December 2024, 103031-103063.
https://doi.org/10.52202/079017-3273
[8] Wang, Z., Liu, L., Wang, L. and Zhou, L. (2023) R2GenGPT: Radiology Report Generation with Frozen LLMs. Meta-Radiology, 1, Article ID: 100033.
https://doi.org/10.1016/j.metrad.2023.100033
[9] Liu, C., Tian, Y., Chen, W., Song, Y. and Zhang, Y. (2024) Bootstrapping Large Language Models for Radiology Report Generation. Proceedings of the AAAI Conference on Artificial Intelligence, 38, 18635-18643.
https://doi.org/10.1609/aaai.v38i17.29826
[10] Tanida, T., Müller, P., Kaissis, G. and Rueckert, D. (2023) Interactive and Explainable Region-Guided Radiology Report Generation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-21 June 2023, 7433-7442.
https://doi.org/10.1109/cvpr52729.2023.00718
[11] Vedantam, R., Zitnick, C.L. and Parikh, D. (2015) CIDEr: Consensus-Based Image Description Evaluation. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, 7-12 June 2015, 4566-4575.
https://doi.org/10.1109/cvpr.2015.7299087
[12] Jain, S., Agrawal, A., Saporta, A., et al. (2021) RadGraph: Extracting Clinical Entities and Relations from Radiology Reports. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 6-11 June 2021, 3172-3183.
[13] Chen, C., Li, O., Tao, D., et al. (2019) This Looks like That: Deep Learning for Interpretable Image Recognition. Advances in Neural Information Processing Systems, Vancouver, 8-14 December 2019, 8930-8941.
[14] Kim, E., Kim, S., Seo, M. and Yoon, S. (2021) XProtoNet: Diagnosis in Chest Radiography with Global and Local Explanations. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 20-25 June 2021, 15719-15728.
https://doi.org/10.1109/cvpr46437.2021.01546
[15] Snell, J., Swersky, K. and Zemel, R. (2017) Prototypical Networks for Few-Shot Learning. Advances in Neural In-formation Processing Systems, Long Beach, 4-9 December 2017, 4077-4087.
[16] Chen, T., Kornblith, S., Norouzi, M., et al. (2020) A Simple Framework for Contrastive Learning of Visual Representations. International Conference on Machine Learning, 13-18 July 2020, 1597-1607.
[17] He, K., Fan, H., Wu, Y., Xie, S. and Girshick, R. (2020) Momentum Contrast for Unsupervised Visual Representation Learning. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, 13-19 June 2020, 9726-9735.
https://doi.org/10.1109/cvpr42600.2020.00975
[18] Li, M., Lin, B., Chen, Z., Lin, H., Liang, X. and Chang, X. (2023) Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report Generation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-21 June 2023, 3334-3343.
https://doi.org/10.1109/cvpr52729.2023.00325
[19] Zhang, Z. and Jiang, A. (2024) Interactive Dual-Stream Contrastive Learning for Radiology Report Generation. Journal of Biomedical Informatics, 157, Article ID: 104718.
https://doi.org/10.1016/j.jbi.2024.104718
[20] Gu, A. and Dao, T. (2023) Mamba: Linear-Time Sequence Modeling with Selective State Spaces.
https://arxiv.org/pdf/2312.00752
[21] Touvron, H., Martin, L., Stone, K., et al. (2023) Llama 2: Open Foundation and Fine-Tuned Chat Models.
https://arxiv.org/abs/2307.09288
[22] Hu, E.J., Shen, Y., Wallis, P., et al. (2022) LoRA: Low-Rank Adaptation of Large Language Models. International Conference on Learning Representations, 25-29 April 2022.
https://openreview.net/forum?id=nZeVKeeFYf9
[23] Zhu, L., Liao, B., Zhang, Q., et al. (2024) Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. Proceedings of the 41st International Conference on Machine Learning, Vienna, 21-27 July 2024, 62429-62442.
[24] Azad, R., Kazerouni, A., Heidari, M., Aghdam, E.K., Molaei, A., Jia, Y., et al. (2024) Advances in Medical Image Analysis with Vision Transformers: A Comprehensive Review. Medical Image Analysis, 91, Article ID: 103000.
https://doi.org/10.1016/j.media.2023.103000
[25] Lin, C.Y. (2004) ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out, Barcelona, 25-26 July 2004, 74-81.
[26] Banerjee, S. and Lavie, A. (2005) METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Ann Arbor, 29 June 2005, 65-72.
[27] Delbrouck, J., Chambon, P., Bluethgen, C., Tsai, E., Almusa, O. and Langlotz, C. (2022) Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards. Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, 7-11 December 2022, 4348-4360.
https://doi.org/10.18653/v1/2022.findings-emnlp.319
[28] Park, S., Kim, J., Lee, H., et al. (2025) KIA: Knowledge-Infused Attention for Accurate Radiology Report Generation. Proceedings of the 31st International Conference on Computational Linguistics (COLING), Abu Dhabi, 19-24 January 2025, 1-12.
[29] Yang, Y., Yu, J., Fu, Z., Zhang, K., Yu, T., Wang, X., et al. (2024) Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image Reporting. IEEE Transactions on Medical Imaging, 43, 4017-4028.
https://doi.org/10.1109/tmi.2024.3412402
[30] Liu, F., Wu, X., Ge, S., Fan, W. and Zou, Y. (2021) Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 20-25 June 2021, 13753-13762.
https://doi.org/10.1109/cvpr46437.2021.01354
[31] Wang, Z., Liu, L., Wang, L. and Zhou, L. (2023) METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert Tokens. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, 17-21 June 2023, 11558-11567.
https://doi.org/10.1109/cvpr52729.2023.01112
[32] Liu, A., Guo, Y., Yong, J-H. and Xu, F. (2024) Multi-Grained Radiology Report Generation with Sentence-Level Image-Language Contrastive Learning. IEEE Transactions on Medical Imaging, 43, 2657-2669.
https://doi.org/10.1109/tmi.2024.3372638