IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v113y2017icp475-496.html
   My bibliography  Save this article

Model-based time-varying clustering of multivariate longitudinal data with covariates and outliers

Author

Listed:
  • Maruotti, Antonello
  • Punzo, Antonio

Abstract

A class of multivariate linear models under the longitudinal setting, in which unobserved heterogeneity may evolve over time, is introduced. A latent structure is considered to model heterogeneity, having a discrete support and following a first-order Markov chain. Heavy-tailed multivariate distributions are introduced to deal with outliers. Maximum likelihood estimation is performed to estimate parameters by using expectation–maximization and expectation–conditional-maximization algorithms. Notes on model identifiability and robustness are provided, along with all computational details needed to implement the proposal. Three applications on artificial and real data are illustrated. These focus on the potential effects of outliers on clustering and their identification.

Suggested Citation

  • Maruotti, Antonello & Punzo, Antonio, 2017. "Model-based time-varying clustering of multivariate longitudinal data with covariates and outliers," Computational Statistics & Data Analysis, Elsevier, vol. 113(C), pages 475-496.
  • Handle: RePEc:eee:csdana:v:113:y:2017:i:c:p:475-496
    DOI: 10.1016/j.csda.2016.05.024
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167947316301360
    Download Restriction: Full text for ScienceDirect subscribers only.

    File URL: https://libkey.io/10.1016/j.csda.2016.05.024?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Maia Berkane & P. M. Bentler, 1988. "Estimation of Contamination Parameters and Identification of Outliers in Multivariate Data," Sociological Methods & Research, , vol. 17(1), pages 55-64, August.
    2. Leroux, Brian G., 1992. "Maximum-likelihood estimation for hidden Markov models," Stochastic Processes and their Applications, Elsevier, vol. 40(1), pages 127-143, February.
    3. Antonello Maruotti, 2011. "Mixed Hidden Markov Models for Longitudinal Data: An Overview," International Statistical Review, International Statistical Institute, vol. 79(3), pages 427-454, December.
    4. Sanjeena Subedi & Antonio Punzo & Salvatore Ingrassia & Paul McNicholas, 2015. "Cluster-weighted $$t$$ t -factor analyzers for robust model-based clustering and dimension reduction," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 24(4), pages 623-649, November.
    5. Hajo Holzmann & Axel Munk & Tilmann Gneiting, 2006. "Identifiability of Finite Mixtures of Elliptical Distributions," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 33(4), pages 753-763, December.
    6. Juárez, Miguel A. & Steel, Mark F. J., 2010. "Model-Based Clustering of Non-Gaussian Panel Data Based on Skew-t Distributions," Journal of Business & Economic Statistics, American Statistical Association, vol. 28(1), pages 52-66.
    7. Ingrassia, Salvatore & Minotti, Simona C. & Punzo, Antonio, 2014. "Model-based clustering via linear cluster-weighted models," Computational Statistics & Data Analysis, Elsevier, vol. 71(C), pages 159-182.
    8. Luca Bagnato & Antonio Punzo, 2013. "Finite mixtures of unimodal beta and gamma densities and the $$k$$ -bumps algorithm," Computational Statistics, Springer, vol. 28(4), pages 1571-1597, August.
    9. Francesca Greselin & Antonio Punzo, 2013. "Closed Likelihood Ratio Testing Procedures to Assess Similarity of Covariance Matrices," The American Statistician, Taylor & Francis Journals, vol. 67(3), pages 117-128, August.
    10. Jesse D. Raffa & Joel A. Dubin, 2015. "Multivariate longitudinal data analysis with mixed effects hidden Markov models," Biometrics, The International Biometric Society, vol. 71(3), pages 821-831, September.
    11. Salvatore Ingrassia & Antonio Punzo & Giorgio Vittadini & Simona Minotti, 2015. "The Generalized Linear Mixed Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 32(1), pages 85-113, April.
    12. Sanjeena Subedi & Antonio Punzo & Salvatore Ingrassia & Paul McNicholas, 2013. "Clustering and classification via cluster-weighted factor analyzers," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 7(1), pages 5-40, March.
    13. Luis García-Escudero & Alfonso Gordaliza & Carlos Matrán & Agustín Mayo-Iscar, 2010. "A review of robust clustering methods," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 4(2), pages 89-109, September.
    14. Goldfeld, Stephen M. & Quandt, Richard E., 1973. "A Markov model for switching regressions," Journal of Econometrics, Elsevier, vol. 1(1), pages 3-15, March.
    15. Francesca Greselin & Salvatore Ingrassia & Antonio Punzo, 2011. "Assessing the pattern of covariance matrices via an augmentation multiple testing procedure," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 20(2), pages 141-170, June.
    16. Roderick J. A. Little, 1988. "Robust Estimation of the Mean and Covariance Matrix from Data with Missing Values," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 37(1), pages 23-38, March.
    17. Biernacki, Christophe & Celeux, Gilles & Govaert, Gerard, 2003. "Choosing starting values for the EM algorithm for getting the highest likelihood in multivariate Gaussian mixture models," Computational Statistics & Data Analysis, Elsevier, vol. 41(3-4), pages 561-575, January.
    18. Salvatore Ingrassia & Antonio Punzo & Giorgio Vittadini & Simona Minotti, 2015. "Erratum to: The Generalized Linear Mixed Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 32(2), pages 327-355, July.
    19. Francesco Lagona & Antonello Maruotti & Fabio Padovano, 2015. "Multilevel multivariate modelling of legislative count data, with a hidden Markov chain," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 178(3), pages 705-723, June.
    20. Turner, Rolf, 2008. "Direct maximization of the likelihood of a hidden Markov model," Computational Statistics & Data Analysis, Elsevier, vol. 52(9), pages 4147-4160, May.
    21. Fruhwirth-Schnatter, Sylvia & Kaufmann, Sylvia, 2008. "Model-Based Clustering of Multiple Time Series," Journal of Business & Economic Statistics, American Statistical Association, vol. 26, pages 78-89, January.
    22. Hamilton, James D., 1990. "Analysis of time series subject to changes in regime," Journal of Econometrics, Elsevier, vol. 45(1-2), pages 39-70.
    23. Sharon Lee & Geoffrey McLachlan, 2013. "Rejoinder to the discussion of “Model-based clustering and classification with non-normal mixture distributions”," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 22(4), pages 473-479, November.
    24. García-Escudero, L.A. & Gordaliza, A. & Mayo-Iscar, A. & San Martín, R., 2010. "Robust clusterwise linear regression through trimming," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3057-3069, December.
    25. Bai, Xiuqin & Yao, Weixin & Boyer, John E., 2012. "Robust fitting of mixture regression models," Computational Statistics & Data Analysis, Elsevier, vol. 56(7), pages 2347-2359.
    26. Iain L. MacDonald, 2014. "Numerical Maximisation of Likelihood: A Neglected Alternative to EM?," International Statistical Review, International Statistical Institute, vol. 82(2), pages 296-308, August.
    27. Inmaculada Martinez‐Zarzoso & Antonello Maruotti, 2013. "The environmental Kuznets curve: functional form, time‐varying heterogeneity and outliers in a panel setting," Environmetrics, John Wiley & Sons, Ltd., vol. 24(7), pages 461-475, November.
    28. Sharon Lee & Geoffrey McLachlan, 2013. "Model-based clustering and classification with non-normal mixture distributions," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 22(4), pages 427-454, November.
    29. Francesco Bartolucci & Alessio Farcomeni, 2015. "A discrete time event-history approach to informative drop-out in mixed latent Markov models with covariates," Biometrics, The International Biometric Society, vol. 71(1), pages 80-89, March.
    30. Bartolucci, Francesco & Farcomeni, Alessio, 2009. "A Multivariate Extension of the Dynamic Logit Model for Longitudinal Data Based on a Latent Markov Heterogeneity Structure," Journal of the American Statistical Association, American Statistical Association, vol. 104(486), pages 816-831.
    31. Alessio Farcomeni & Luca Greco, 2015. "S-estimation of hidden Markov models," Computational Statistics, Springer, vol. 30(1), pages 57-80, March.
    32. Jan Bulla & Andreas Berzel, 2008. "Computational issues in parameter estimation for stationary hidden Markov models," Computational Statistics, Springer, vol. 23(1), pages 1-18, January.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Antonio Punzo & Paul. D. McNicholas, 2017. "Robust Clustering in Regression Analysis via the Contaminated Gaussian Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 34(2), pages 249-293, July.
    2. Cremaschini, Alessandro & Maruotti, Antonello, 2023. "A finite mixture analysis of structural breaks in the G-7 gross domestic product series," Research in Economics, Elsevier, vol. 77(1), pages 76-90.
    3. Bucci, Alberto & Carbonari, Lorenzo & Gil, Pedro Mazeda & Trovato, Giovanni, 2021. "Economic growth and innovation complexity: An empirical estimation of a Hidden Markov Model," Economic Modelling, Elsevier, vol. 98(C), pages 86-99.
    4. Angelo Mazza & Antonio Punzo, 2020. "Mixtures of multivariate contaminated normal regression models," Statistical Papers, Springer, vol. 61(2), pages 787-822, April.
    5. Wan-Lun Wang, 2019. "Mixture of multivariate t nonlinear mixed models for multiple longitudinal data with heterogeneity and missing values," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(1), pages 196-222, March.
    6. Morris, Katherine & Punzo, Antonio & McNicholas, Paul D. & Browne, Ryan P., 2019. "Asymmetric clusters and outliers: Mixtures of multivariate contaminated shifted asymmetric Laplace distributions," Computational Statistics & Data Analysis, Elsevier, vol. 132(C), pages 145-166.
    7. Antonio Punzo & Salvatore Ingrassia & Antonello Maruotti, 2021. "Multivariate hidden Markov regression models: random covariates and heavy-tailed distributions," Statistical Papers, Springer, vol. 62(3), pages 1519-1555, June.
    8. Benjamin Auder & Elisabeth Gassiat & Mor Absa Loum, 2021. "Least squares moment identification of binary regression mixture models," Metrika: International Journal for Theoretical and Applied Statistics, Springer, vol. 84(4), pages 561-593, May.
    9. Yang, Yu-Chen & Lin, Tsung-I & Castro, Luis M. & Wang, Wan-Lun, 2020. "Extending finite mixtures of t linear mixed-effects models with concomitant covariates," Computational Statistics & Data Analysis, Elsevier, vol. 148(C).
    10. Antonello Maruotti & Antonio Punzo, 2021. "Initialization of Hidden Markov and Semi‐Markov Models: A Critical Evaluation of Several Strategies," International Statistical Review, International Statistical Institute, vol. 89(3), pages 447-480, December.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Angelo Mazza & Antonio Punzo, 2020. "Mixtures of multivariate contaminated normal regression models," Statistical Papers, Springer, vol. 61(2), pages 787-822, April.
    2. Antonio Punzo & Paul. D. McNicholas, 2017. "Robust Clustering in Regression Analysis via the Contaminated Gaussian Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 34(2), pages 249-293, July.
    3. Salvatore Ingrassia & Antonio Punzo, 2020. "Cluster Validation for Mixtures of Regressions via the Total Sum of Squares Decomposition," Journal of Classification, Springer;The Classification Society, vol. 37(2), pages 526-547, July.
    4. Antonio Punzo & Salvatore Ingrassia & Antonello Maruotti, 2021. "Multivariate hidden Markov regression models: random covariates and heavy-tailed distributions," Statistical Papers, Springer, vol. 62(3), pages 1519-1555, June.
    5. Antonello Maruotti & Antonio Punzo, 2021. "Initialization of Hidden Markov and Semi‐Markov Models: A Critical Evaluation of Several Strategies," International Statistical Review, International Statistical Institute, vol. 89(3), pages 447-480, December.
    6. Yang, Yu-Chen & Lin, Tsung-I & Castro, Luis M. & Wang, Wan-Lun, 2020. "Extending finite mixtures of t linear mixed-effects models with concomitant covariates," Computational Statistics & Data Analysis, Elsevier, vol. 148(C).
    7. Paul D. McNicholas, 2016. "Model-Based Clustering," Journal of Classification, Springer;The Classification Society, vol. 33(3), pages 331-373, October.
    8. Salvatore Ingrassia & Antonio Punzo & Giorgio Vittadini & Simona Minotti, 2015. "Erratum to: The Generalized Linear Mixed Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 32(2), pages 327-355, July.
    9. Diani, Cecilia & Galimberti, Giuliano & Soffritti, Gabriele, 2022. "Multivariate cluster-weighted models based on seemingly unrelated linear regression," Computational Statistics & Data Analysis, Elsevier, vol. 171(C).
    10. Maruotti, Antonello & Petrella, Lea & Sposito, Luca, 2021. "Hidden semi-Markov-switching quantile regression for time series," Computational Statistics & Data Analysis, Elsevier, vol. 159(C).
    11. Morris, Katherine & Punzo, Antonio & McNicholas, Paul D. & Browne, Ryan P., 2019. "Asymmetric clusters and outliers: Mixtures of multivariate contaminated shifted asymmetric Laplace distributions," Computational Statistics & Data Analysis, Elsevier, vol. 132(C), pages 145-166.
    12. Salvatore D. Tomarchio & Luca Bagnato & Antonio Punzo, 2022. "Model-based clustering via new parsimonious mixtures of heavy-tailed distributions," AStA Advances in Statistical Analysis, Springer;German Statistical Society, vol. 106(2), pages 315-347, June.
    13. Počuča, Nikola & Jevtić, Petar & McNicholas, Paul D. & Miljkovic, Tatjana, 2020. "Modeling frequency and severity of claims with the zero-inflated generalized cluster-weighted models," Insurance: Mathematics and Economics, Elsevier, vol. 94(C), pages 79-93.
    14. Gordon Anderson & Alessio Farcomeni & Maria Grazia Pittau & Roberto Zelli, 2019. "Rectangular latent Markov models for time‐specific clustering, with an analysis of the wellbeing of nations," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 68(3), pages 603-621, April.
    15. Michael P. B. Gallaugher & Salvatore D. Tomarchio & Paul D. McNicholas & Antonio Punzo, 2022. "Multivariate cluster weighted models using skewed distributions," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 16(1), pages 93-124, March.
    16. Utkarsh J. Dang & Antonio Punzo & Paul D. McNicholas & Salvatore Ingrassia & Ryan P. Browne, 2017. "Multivariate Response and Parsimony for Gaussian Cluster-Weighted Models," Journal of Classification, Springer;The Classification Society, vol. 34(1), pages 4-34, April.
    17. Salvatore Ingrassia & Antonio Punzo & Giorgio Vittadini & Simona Minotti, 2015. "The Generalized Linear Mixed Cluster-Weighted Model," Journal of Classification, Springer;The Classification Society, vol. 32(1), pages 85-113, April.
    18. Naderi, Mehrdad & Mirfarah, Elham & Wang, Wan-Lun & Lin, Tsung-I, 2023. "Robust mixture regression modeling based on the normal mean-variance mixture distributions," Computational Statistics & Data Analysis, Elsevier, vol. 180(C).
    19. Faicel Chamroukhi, 2016. "Piecewise Regression Mixture for Simultaneous Functional Data Clustering and Optimal Segmentation," Journal of Classification, Springer;The Classification Society, vol. 33(3), pages 374-411, October.
    20. Paolo Berta & Salvatore Ingrassia & Antonio Punzo & Giorgio Vittadini, 2016. "Multilevel cluster-weighted models for the evaluation of hospitals," METRON, Springer;Sapienza Università di Roma, vol. 74(3), pages 275-292, December.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:113:y:2017:i:c:p:475-496. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.