IDEAS home Printed from https://ideas.repec.org/p/ems/eureir/77010.html
   My bibliography  Save this paper

Cluster Correspondence Analysis

Author

Listed:
  • van de Velden, M.
  • Iodice D' Enza, A.
  • Palumbo, F.

Abstract

__Abstract__ A new method is proposed that combines dimension reduction and cluster analysis for categorical data. A least-squares objective function is formulated that approximates the cluster by variables cross-tabulation. Individual observations are assigned to clusters in such a way that the distributions over the categorical variables for the different clusters are optimally separated. In a unified framework, a brief review of alternative methods is provided and performance of the methods is appraised by means of a simulation study. The results of the joint dimension reduction and clustering methods are compared with cluster analysis based on the full dimensional data. Our results show that the joint dimension reduction and clustering methods outperform, both with respect to the retrieval of the true underlying cluster structure and with respect to internal cluster validity measures, full dimensional clustering. The differences increase when more variables are involved and in the presence of noise variables.

Suggested Citation

  • van de Velden, M. & Iodice D' Enza, A. & Palumbo, F., 2014. "Cluster Correspondence Analysis," Econometric Institute Research Papers EI 2014-24, Erasmus University Rotterdam, Erasmus School of Economics (ESE), Econometric Institute.
  • Handle: RePEc:ems:eureir:77010
    as

    Download full text from publisher

    File URL: https://repub.eur.nl/pub/77010/EI2014-24.pdf
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
    2. Vichi, Maurizio & Kiers, Henk A. L., 2001. "Factorial k-means analysis for two-way data," Computational Statistics & Data Analysis, Elsevier, vol. 37(1), pages 49-64, July.
    3. Michel Velden & Yoshio Takane, 2012. "Generalized canonical correlation analysis with missing values," Computational Statistics, Springer, vol. 27(3), pages 551-571, September.
    4. Alfonso Iodice D’Enza & Francesco Palumbo, 2013. "Iterative factor clustering of binary data," Computational Statistics, Springer, vol. 28(2), pages 789-807, April.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Ming Sun & Ronggui Zhou, 2023. "Investigation on Hazardous Material Truck Involved Fatal Crashes Using Cluster Correspondence Analysis," Sustainability, MDPI, vol. 15(12), pages 1-21, June.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. M. Velden & A. Iodice D’Enza & F. Palumbo, 2017. "Cluster Correspondence Analysis," Psychometrika, Springer;The Psychometric Society, vol. 82(1), pages 158-185, March.
    2. Masaki Mitsuhiro & Hiroshi Yadohisa, 2015. "Reduced $$k$$ k -means clustering with MCA in a low-dimensional space," Computational Statistics, Springer, vol. 30(2), pages 463-475, June.
    3. DeSarbo, Wayne S. & Selin Atalay, A. & Blanchard, Simon J., 2009. "A three-way clusterwise multidimensional unfolding procedure for the spatial representation of context dependent preferences," Computational Statistics & Data Analysis, Elsevier, vol. 53(8), pages 3217-3230, June.
    4. Roberto Rocci & Stefano Gattone & Maurizio Vichi, 2011. "A New Dimension Reduction Method: Factor Discriminant K-means," Journal of Classification, Springer;The Classification Society, vol. 28(2), pages 210-226, July.
    5. Naoto Yamashita & Shin-ichi Mayekawa, 2015. "A new biplot procedure with joint classification of objects and variables by fuzzy c-means clustering," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 9(3), pages 243-266, September.
    6. Donatella Vicari & Paolo Giordani, 2023. "CPclus: Candecomp/Parafac Clustering Model for Three-Way Data," Journal of Classification, Springer;The Classification Society, vol. 40(2), pages 432-465, July.
    7. Michael C. Thrun & Alfred Ultsch, 2021. "Using Projection-Based Clustering to Find Distance- and Density-Based Clusters in High-Dimensional Data," Journal of Classification, Springer;The Classification Society, vol. 38(2), pages 280-312, July.
    8. Michio Yamamoto, 2012. "Clustering of functional data in a low-dimensional subspace," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 6(3), pages 219-247, October.
    9. Cristina Tortora & Paul D. McNicholas & Ryan P. Browne, 2016. "A mixture of generalized hyperbolic factor analyzers," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 10(4), pages 423-440, December.
    10. Efthymios Costa & Ioanna Papatsouma & Angelos Markos, 2023. "Benchmarking distance-based partitioning methods for mixed-type data," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 17(3), pages 701-724, September.
    11. Kensuke Tanioka & Hiroshi Yadohisa, 2019. "Simultaneous Method of Orthogonal Non-metric Non-negative Matrix Factorization and Constrained Non-hierarchical Clustering," Journal of Classification, Springer;The Classification Society, vol. 36(1), pages 73-93, April.
    12. Monia Ranalli & Roberto Rocci, 2017. "A Model-Based Approach to Simultaneous Clustering and Dimensional Reduction of Ordinal Data," Psychometrika, Springer;The Psychometric Society, vol. 82(4), pages 1007-1034, December.
    13. Timmerman, Marieke E. & Ceulemans, Eva & Kiers, Henk A.L. & Vichi, Maurizio, 2010. "Factorial and reduced K-means reconsidered," Computational Statistics & Data Analysis, Elsevier, vol. 54(7), pages 1858-1871, July.
    14. Yoshikazu Terada, 2014. "Strong Consistency of Reduced K-means Clustering," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 41(4), pages 913-931, December.
    15. Luca Greco & Antonio Lucadamo & Pietro Amenta, 2020. "An Impartial Trimming Approach for Joint Dimension and Sample Reduction," Journal of Classification, Springer;The Classification Society, vol. 37(3), pages 769-788, October.
    16. Fordellone, Mario & Vichi, Maurizio, 2020. "Finding groups in structural equation modeling through the partial least squares algorithm," Computational Statistics & Data Analysis, Elsevier, vol. 147(C).
    17. Michio Yamamoto & Heungsun Hwang, 2017. "Dimension-Reduced Clustering of Functional Data via Subspace Separation," Journal of Classification, Springer;The Classification Society, vol. 34(2), pages 294-326, July.
    18. Alicja Grześkowiak, 2016. "Assessment of Participation in Cultural Activities in Poland by Selected Multivariate Methods," European Journal of Social Sciences Education and Research Articles, Revistia Research and Publishing, vol. 3, January -.
    19. Yunpeng Zhao & Qing Pan & Chengan Du, 2019. "Logistic regression augmented community detection for network data with application in identifying autism‐related gene pathways," Biometrics, The International Biometric Society, vol. 75(1), pages 222-234, March.
    20. Wu, Han-Ming & Tien, Yin-Jing & Chen, Chun-houh, 2010. "GAP: A graphical environment for matrix visualization and cluster analysis," Computational Statistics & Data Analysis, Elsevier, vol. 54(3), pages 767-778, March.

    More about this item

    Keywords

    Correspondence analysis; cluster analysis; dimension; reduction; categorical variables;
    All these keywords.

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:ems:eureir:77010. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: RePub (email available below). General contact details of provider: https://edirc.repec.org/data/feeurnl.html .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.