面向形态多变垃圾的轻量化视觉语言模型细粒度分类方法研究
Research on Fine-Grained Classification Method of Lightweight Vision-Language Model for Morphologically Variable Waste
摘要: 针对垃圾分类场景下垃圾形态多变、细粒度类别特征相近、训练样本稀缺以及边缘部署算力不足等难题,本文基于轻量化视觉语言模型SmolVLM-256,开展小样本细粒度垃圾分类的性能优化研究。结合真实场景中垃圾挤压、遮挡、复杂背景干扰等特点,本文构建了包含8类生活垃圾的细粒度数据集并搭建标准化实验评估框架。通过超参数调优、语义提示引导、数据增强与分类头结构改良等多维度优化实验,探究各策略对模型分类效果的影响。实验结果显示,优化后的SmolVLM-256分类准确率可达87.92% ± 1.56%,相较基线模型性能大幅提升。横向对比ViT、CLIP、BLIP、Qwen3-VL等模型可知,该轻量化模型在保障优良分类精度的同时,具备参数量小、推理成本低的优势,更适配边缘设备部署。研究表明,合理的训练与结构优化策略,能够有效提升轻量化视觉语言模型在复杂小样本垃圾场景的识别能力,可为边缘智能垃圾分类系统的研发与落地提供技术参考。
Abstract: Aiming at the practical challenges of variable garbage morphology, similar fine-grained feature distribution, limited training samples, and insufficient computing resources for edge deployment in garbage classification scenarios, this paper conducts a performance optimization study on small-sample fine-grained garbage classification based on the lightweight vision-language model SmolVLM-256. Considering real-world interferences such as garbage extrusion, occlusion, and complex background environments, a fine-grained dataset containing eight types of domestic garbage is constructed, and a standardized experimental evaluation framework is established. Multiple optimization experiments including hyperparameter tuning, semantic prompt guidance, data augmentation, and classification head structure improvement are carried out to explore the influence of different strategies on model classification performance. The experimental results show that the optimized SmolVLM-256 achieves a classification accuracy of 87.92% ± 1.56%, which obtains a significant performance improvement compared with the baseline model. Comparative experiments with ViT, CLIP, BLIP, and Qwen3-VL demonstrate that the lightweight model maintains competitive classification accuracy with fewer parameters and lower inference cost, making it more suitable for edge deployment. The research verifies that reasonable training strategies and structural optimization can effectively enhance the recognition capability of lightweight vision-language models in complex small-sample garbage scenarios, providing technical references for the design and practical application of edge intelligent garbage classification systems.
文章引用:孙紫阳, 刘柱, 王向前, 李娜, 熊旭辉. 面向形态多变垃圾的轻量化视觉语言模型细粒度分类方法研究[J]. 计算机科学与应用, 2026, 16(7): 125-139. https://doi.org/10.12677/csa.2026.167247

参考文献

[1] 肖克江, 陈亮, 高阔, 杨文齐, 庞世燕. 基于多模态数据融合的边缘设备轻量级垃圾分类方法和系统研究[J]. 工程科学学报, 2025, 47(9): 1905-1916.
[2] 魏秀参, 杨晓龙, 李泽超, 等. 细粒度图像分析研究进展与展望[J]. 计算机学报, 2023, 46(4): 861-891.
[3] 王晋东, 陈益强. 小样本学习研究综述[J]. 计算机学报, 2021, 44(11): 2441-2470.
[4] 张正, 王超, 李凯, 等. 深度学习模型轻量化技术综述[J]. 软件学报, 2023, 34(5): 2319-2348.
[5] Marafioti, A., Zohar, O., Farré, M., Noyan, M., Bakouch, E., Cuenca, P., et al. (2025) SmolVLM: Rede-Fining Small and Efficient Multimodal Models.
https://arxiv.org/abs/2504.05299
[6] 刘知远, 孙茂松. 提示学习研究进展与展望[J]. 中国科学: 信息科学, 2023, 53(2): 247-273.
[7] 孙紫阳, 刘柱, 王向前, 李娜. 面向形态多变垃圾的细粒度生活垃圾数据集[EB/OL].
https://github.com/1583118323/Waste_FineGrained_Dataset, 2025-12-20.
[8] Bergstra, J. and Bengio, Y. (2012) Random Search for Hyper-Parameter Optimization. Journal of Machine Learning Re-search, 13, 281-305.
[9] Cao, Y., Liu, Y.Z., Wang, W.H., et al. (2024) MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding. IEEE Transactions on Image Processing, 33, 4721-4734.
[10] Chen, Q.P., Jiao, L., Wang, F.M., Du, J.M., Liu, H.Y., Wang, X. and Wang, R.J. (2024) Integrating Foreground-Background Feature Distillation and Contrastive Feature Learning for Ultra-Fine-Grained Visual Classification. Pattern Recognition, 150, 110339. [Google Scholar] [CrossRef
[11] Shorten, C. and Khoshgoftaar, T.M. (2019) A Survey on Image Data Augmentation for Deep Learning. Journal of Big Data, 6, Article No. 60. [Google Scholar] [CrossRef
[12] 李玉洁, 马子航, 王艺甫, 等. 视觉Transformer (ViT)发展综述[J]. 计算机科学, 2025, 52(1): 194-209.
[13] Vaswani, A., Shazeer, N., Parmar, N., et al. (2017) Attention Is All You Need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, 4-9 December 2017, 6000-6010.
[14] Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2021) An Image Is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations, 3-7 May 2021, 1-16.
[15] Li, J., Li, D., Xiong, C. and Hoi, S. (2022) BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation. Proceedings of the 39th International Conference on Machine Learning, Baltimore, 17-23 July 2022, 12888-12900.
[16] Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al. (2021) Learning Transferable Visual Models from Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning, 18-24 July 2021, 8748-8763.
[17] Bai, J., Bai, Y., Wang, Z., et al. (2023) Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.
https://arxiv.org/abs/2308.12966
[18] 刘群, 李鹏, 王士进, 等. 视觉语言预训练模型研究综述[J]. 计算机学报, 2024, 47(2): 341-372.
[19] 李卫. 视觉Transformer模型设计及其轻量化研究[D]: [硕士学位论文]. 成都: 电子科技大学, 2023.
[20] Srivastava, N., Hinton, G., Krizhevsky, A., et al. (2014) Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research, 15, 1929-1958.
[21] Wu, W., Zhang, Y., Li, X., et al. (2023) Applications of Convolutional Neural Networks for Intelligent Waste Identification and Recycling: A Review. Journal of Cleaner Production, 397, 136542. [Google Scholar] [CrossRef
[22] He, K., Zhang, X., Ren, S. and Sun, J. (2016) Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, 27-30 June 2016, 770-778. [Google Scholar] [CrossRef
[23] Howard, A., Zhu, M., Chen, B., et al. (2017) MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.
https://arxiv.org/abs/1704.04861
[24] Chen, G., Li, Y. and Wang, H. (2025) RepVGG-MEM: A Lightweight Model for Garbage Classification Achieving a Balance between Accuracy and Speed. IEEE Access, 13, 28764-28775. [Google Scholar] [CrossRef
[25] Wang, L., Zhang, Q. and Liu, H. (2023) Integrating Human Vision Perception in Vision Transformers for Classifying Waste Items. Applied Sciences, 13, 13215.
[26] Chen, X., Li, Y., Zhang, H., et al. (2024) Zero-Shot Garbage Classification Using CLIP with Prompt Engineering. 2024 IEEE International Conference on Multimedia and Expo, Niagara Falls, 15-19 July 2024, 1-6.
[27] Zhang, C., Liu, Y. and Wang, H. (2025) Hierarchical Prompt Design for Fine-Grained Image Classification. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, 11-15 June 2025, 123-130.
[28] 黄凯奇, 陈晓棠, 康运锋, 等. 基于深度学习的垃圾分类方法研究进展[J]. 自动化学报, 2022, 48(3): 647-666.
[29] Hu, E.J., Shen, Y., Wallis, P., et al. (2022) LoRA: Low-Rank Adaptation of Large Language Models. International Conference on Learning Representations, 25-29 April 2022, 1-17.
[30] Shin, T., Razeghi, Y., Logan IV, R.L., Wallace, E. and Singh, S. (2020) AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 16-20 November 2020, 4222-4235. [Google Scholar] [CrossRef
[31] Frantar, E., Ashkboos, S., Hoefler, T., et al. (2023) GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers.
https://arxiv.org/abs/2210.17323
[32] Hinton, G., Vinyals, O. and Dean, J. (2015) Distilling the Knowledge in a Neural Network.
https://arxiv.org/abs/1503.02531
[33] Wang, Z., Li, Y. and Zhang, H. (2025) Continual Learning for Vision-Language Models: A Survey.
https://arxiv.org/abs/2501.07654