基于有限混合回归模型的各国基尼指数纵向聚类研究
Longitudinal Clustering Analysis of National Gini Coefficients Based on Finite Mixture Regression Model
摘要: 纵向数据的异质性特征在经济学与社会科学领域十分常见,比如不同国家收入不平等的演化轨迹具有明显异质性,对其进行聚类分析有助于揭示不同国家经济发展的潜在模式与异质性规律。在处理纵向数据的聚类分析中,有限混合回归模型(Finite Mixture Regression Model, FMR)是一类核心工具。现有的混合建模主要包含两种思路:传统有限混合回归模型(FMR-1)通常假设具有个体内部观测值相互独立,仅对各子总体的均值进行建模。然而,这种忽略组内协方差信息的做法易导致聚类偏差。针对此局限,有学者提出了基于均值–协方差联合建模的混合回归模型(FMR-3),其在正态混合框架下同时对均值与协方差结构进行建模。本文运用上述两种方法对36个国家20年间的基尼系数演化轨迹进行聚类分析,两者均采用EM算法进行参数估计。结果表明,基于均值–协方差联合建模的混合回归方法在聚类一致性、组内轨迹区分度以及模型拟合优度(BIC)上均显著优于传统均值模型;而忽略组内相关性不仅容易导致聚类数目的高估,还会造成不同演化轨迹的相互混杂。本研究证实了在实际宏观经济纵向数据聚类中纳入协方差结构的必要性,为跨国收入不平等的动态演化研究提供了更稳健的分析工具。
Abstract: Heterogeneity is a pervasive characteristic of longitudinal data in economics and social sciences. Specifically, the evolutionary trajectories of income inequality display pronounced cross-national heterogeneity, and cluster analysis of such trajectories facilitates the identification of underlying patterns of economic development and heterogeneous dynamics among nations. As a core instrument for longitudinal data clustering, finite mixture regression (FMR) models are widely adopted in the literature. Existing mixture modeling frameworks fall into two main categories: the conventional finite mixture regression model (denoted FMR-1) generally assumes independence among observations sharing the same trajectory pattern and only specifies the mean structure for each sub-population. Nevertheless, overlooking within-group covariance information in this manner is prone to clustering bias. To address this limitation, a mixture regression model with joint mean-covariance modeling (denoted FMR-3) has been developed, which jointly models both the mean and covariance structures within the finite normal mixture framework. Accordingly, this paper employs both approaches to conduct cluster analysis on the Gini coefficient trajectories of 36 countries over a 20-year horizon, with parameter estimation for both models implemented via the Expectation-Maximization (EM) algorithm. Empirical results indicate that the joint mean-covariance mixture regression model substantially outperforms the conventional mean-only model across three dimensions: clustering consistency, within-group trajectory discriminability, and model goodness-of-fit as measured by the Bayesian Information Criterion (BIC). Meanwhile, neglecting within-group correlation not only tends to overestimate the optimal number of clusters, but also causes the conflation of distinct evolutionary trajectories. This study corroborates the necessity of incorporating covariance structures when clustering real-world macroeconomic longitudinal data, and provides a more robust analytical framework for investigating the dynamic evolution of cross-national income inequality.
文章引用:史丁伊, 余静. 基于有限混合回归模型的各国基尼指数纵向聚类研究[J]. 统计学与应用, 2026, 15(8): 1-10. https://doi.org/10.12677/sa.2026.158174

参考文献

[1] Espoir, D.K. (2022) Convergence or Divergence Patterns in Income Distribution across Countries: A New Evidence from a Club Clustering Algorithm. Cogent Economics & Finance, 10, Article 2025667.
https://doi.org/10.1080/23322039.2022.2025667
[2] Ogundari, K. (2023) Club Convergence in Income Inequality in Africa. Social Indicators Research, 167, 319-337.
https://doi.org/10.1007/s11205-023-03108-7
[3] Chrisendo, D., Niva, V., Hoffmann, R., Masoumzadeh Sayyar, S., Rocha, J., Sandström, V., et al. (2025) Rising Income Inequality across Half of Global Population and Socioecological Implications. Nature Sustainability, 8, 1601-1613.
https://doi.org/10.1038/s41893-025-01689-4
[4] McLachlan, G. and Peel, D. (2000) Finite Mixture Models. Wiley.
https://doi.org/10.1002/0471721182
[5] Bouguila, N. and Fan, W. (2020) Mixture Models and Applications. Springer.
[6] Zamani, H., Faroughi, P. and Ismail, N. (2014) Estimation of Count Data Using Mixed Poisson, Generalized Poisson and Finite Poisson Mixture Regression Models. AIP Conference Proceedings, Kuala Lumpur, 17-19 December 2013, 1144-1150.
https://doi.org/10.1063/1.4882628
[7] Nagin, D.S. and Odgers, C.L. (2010) Group-Based Trajectory Modeling in Clinical Research. Annual Review of Clinical Psychology, 6, 109-138.
https://doi.org/10.1146/annurev.clinpsy.121208.131413
[8] Lyrvall, J., Di Mari, R., Bakk, Z., Oser, J. and Kuha, J. (2025) Multilevel Latent Class Analysis: State-of-the-Art Methodologies and Their Implementation in the R Package multilevLCA. Multivariate Behavioral Research, 60, 731-747.
https://doi.org/10.1080/00273171.2025.2473935
[9] Proust-Lima, C., Philipps, V. and Liquet, B. (2017) Estimation of Extended Mixed Models Using Latent Classes and Latent Processes: The R Package lcmm. Journal of Statistical Software, 78, 1-56.
https://doi.org/10.18637/jss.v078.i02
[10] Van der Nest, G., Lima Passos, V., Candel, M.J. and van Breukelen, G.J. (2020) An Overview of Mixture Models for Clustering Longitudinal Data. Advances in Data Analysis and Classification, 14, 361-384.
[11] Wu, L., Li, S. and Tao, Y. (2020) Estimation and Variable Selection for Mixture of Joint Mean and Variance Models. Communications in Statistics-Theory and Methods, 49, 3904-3925.
[12] Yu, J., Nummi, T. and Pan, J. (2022) Mixture Regression for Longitudinal Data Based on Joint Mean-Covariance Model. Journal of Multivariate Analysis, 190, Article 104956.
https://doi.org/10.1016/j.jmva.2022.104956
[13] Yu, J. and Pan, J. (2025) Variable Selection in Mixture Regression for Longitudinal Data Based on Joint Mean-Covariance Model. Journal of Multivariate Analysis, 212, Article 105548.
https://doi.org/10.1016/j.jmva.2025.105548
[14] De la Cruz-Mesía, R., Quintana, F.A. and Marshall, G. (2007) Model-Based Clustering for Longitudinal Data. Computational Statistics & Data Analysis, 51, 6443-6457.
[15] Pourahmadi, M. (1999) Joint Mean-Covariance Models with Applications to Longitudinal Data: Unconstrained Parameterisation. Biometrika, 86, 677-690.
https://doi.org/10.1093/biomet/86.3.677
[16] Cepeda, E. and Gamerman, D. (2001) Bayesian Modeling of Variance Heterogeneity in Normal Regression Models. Brazilian Journal of Probability and Statistics, 15, 139-153.
[17] Pourahmadi, M. (2000) Maximum Likelihood Estimation of Generalised Linear Models for Multivariate Normal Covariance Matrix. Biometrika, 87, 425-435.
https://doi.org/10.1093/biomet/87.2.425