高维小样本下多源财税数据融合的税务风险识别模型研究
Research on a Tax Risk Identification Model Based on Multi-Source Fiscal and Tax Data Fusion in High-Dimensional Small-Sample Scenarios
摘要: 针对涉税数据多源异构、高维稀疏、风险样本稀缺及类别不均衡等问题,本文构建多源财税数据融合的税务风险识别模型。模型分别采用TabNet、轻量化LSTM和图神经网络提取结构化财税、月度经营时序和企业关联拓扑特征,并通过多头注意力机制完成特征融合;同时结合正则化、Dropout、早停和半监督伪标签扩充,以缓解小样本条件下的过拟合。基于匿名化企业涉税样本的实验结果显示,模型AUC达到0.942,F1值较主要基线模型提高12%~25%,说明其在风险样本稀缺场景下具有较好的识别效果和稳定性。研究可为税务风险筛查和企业合规自查提供参考。
Abstract: To address multi-source heterogeneity, high-dimensional sparsity, scarce risk samples, and class imbalance in tax-related data, this paper develops a tax risk identification model based on multi-source fiscal and tax data fusion. TabNet, a lightweight LSTM, and a graph neural network are used to extract structured fiscal and tax features, monthly operational time-series features, and enterprise association topology features, respectively, while a multi-head attention mechanism is introduced for feature fusion. Regularization, Dropout, early stopping, and semi-supervised pseudo-label augmentation are further adopted to alleviate overfitting under small-sample conditions. Experiments on anonymized enterprise tax data show that the model achieves an AUC of 0.942 and improves the F1-score by 12% to 25% over major baseline models, indicating good identification performance and stability when risk samples are scarce. The study provides a reference for tax risk screening and enterprise compliance self-inspection.
文章引用:邹恩源. 高维小样本下多源财税数据融合的税务风险识别模型研究[J]. 应用数学进展, 2026, 15(7): 254-262. https://doi.org/10.12677/aam.2026.157320

参考文献

[1] 李新宇, 朱家明, 徐凤. 基于XGBoost和SHAP方法的中小微企业信贷风险评估研究[J]. 阜阳师范大学学报(自然科学版), 2022, 39(3): 73-81.
[2] 袁涛, 宋加山, 胥兴军, 等. 基于XGBoost算法和SHAP可解释框架的上市公司财务风险预测模型研究[J]. 管理现代化, 2025, 45(6): 69-77.
[3] Hochreiter, S. and Schmidhuber, J. (1997) Long Short-Term Memory. Neural Computation, 9, 1735-1780. [Google Scholar] [CrossRef] [PubMed]
[4] Arik, S.Ö. and Pfister, T. (2021) TabNet: Attentive Interpretable Tabular Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35, 6679-6687. [Google Scholar] [CrossRef
[5] Vaswani, A., Shazeer, N., Parmar, N., et al. (2017) Attention Is All You Need. Advances in Neural Information Processing Systems, Long Beach, 4-9 December 2017, 5998-6008.
[6] Srivastava, N., Hinton, G., Krizhevsky, A., et al. (2014) Dropout: A Simple Way to Prevent Neural Networks from Overfitting. Journal of Machine Learning Research, 15, 1929-1958.
[7] Han, H., Wang, W.Y. and Mao, B.H. (2005) Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning. In: Lecture Notes in Computer Science, Springer, 878-887. [Google Scholar] [CrossRef
[8] Kingma, D.P. and Ba, J. (2014) Adam: A Method for Stochastic Optimization.