An empirical comparison of two approaches for CDPCA in high-dimensional data

My bibliography Save this article

An empirical comparison of two approaches for CDPCA in high-dimensional data

Author

Listed:

Adelaide Freitas
(University of Aveiro
University of Aveiro)
Eloísa Macedo
(University of Aveiro)
Maurizio Vichi
(University “La Sapienza”)

Registered:

Abstract

Modified principal component analysis techniques, specially those yielding sparse solutions, are attractive due to its usefulness for interpretation purposes, in particular, in high-dimensional data sets. Clustering and disjoint principal component analysis (CDPCA) is a constrained PCA that promotes sparsity in the loadings matrix. In particular, CDPCA seeks to describe the data in terms of disjoint (and possibly sparse) components and has, simultaneously, the particularity of identifying clusters of objects. Based on simulated and real gene expression data sets where the number of variables is higher than the number of the objects, we empirically compare the performance of two different heuristic iterative procedures, namely ALS and two-step-SDP algorithms proposed in the specialized literature to perform CDPCA. To avoid possible effect of different variance values among the original variables, all the data was standardized. Although both procedures perform well, numerical tests highlight two main features that distinguish their performance, in particular related to the two-step-SDP algorithm: it provides faster results than ALS and, since it employs a clustering procedure (k-means) on the variables, outperforms ALS algorithm in recovering the true variable partitioning unveiled by the generated data sets. Overall, both procedures produce satisfactory results in terms of solution precision, where ALS performs better, and in recovering the true object clusters, in which two-step-SDP outperforms ALS approach for data sets with lower sample size and more structure complexity (i.e., error level in the CDPCA model). The proportion of explained variance by the components estimated by both algorithms is affected by the data structure complexity (higher error level, the lower variance) and presents similar values for the two algorithms, except for data sets with two object clusters where the two-step-SDP approach yields higher variance. Moreover, experimental tests suggest that the two-step-SDP approach, in general, presents more ability to recover the true number of object clusters, while the ALS algorithm is better in terms of quality of object clustering with more homogeneous, compact and well-separated clusters in the reduced space of the CDPCA components.

Suggested Citation

Adelaide Freitas & Eloísa Macedo & Maurizio Vichi, 2021. "An empirical comparison of two approaches for CDPCA in high-dimensional data," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 30(3), pages 1007-1031, September.

Handle: RePEc:spr:stmapp:v:30:y:2021:i:3:d:10.1007_s10260-020-00546-2
DOI: 10.1007/s10260-020-00546-2

Download full text from publisher

As the access to this document is restricted, you may want to search for a different version of it.

References listed on IDEAS

Carlo Cavicchia & Maurizio Vichi & Giorgia Zaccaria, 2020. "The ultrametric correlation matrix for modelling hierarchical latent concepts," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 14(4), pages 837-853, December.
Rocci, Roberto & Vichi, Maurizio, 2008. "Two-mode multi-partitioning," Computational Statistics & Data Analysis, Elsevier, vol. 52(4), pages 1984-2003, January.
Doyo Enki & Nickolay Trendafilov & Ian Jolliffe, 2013. "A clustering approach to interpretable principal components," Journal of Applied Statistics, Taylor & Francis Journals, vol. 40(3), pages 583-599.
Maurizio Vichi, 2017. "Disjoint factor analysis with cross-loadings," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 11(3), pages 563-591, September.
S. K. Vines, 2000. "Simple principal components," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 49(4), pages 441-451.
Vichi, Maurizio & Saporta, Gilbert, 2009. "Clustering and disjoint principal component analysis," Computational Statistics & Data Analysis, Elsevier, vol. 53(8), pages 3194-3208, June.
Charrad, Malika & Ghazzali, Nadia & Boiteau, Véronique & Niknafs, Azam, 2014. "NbClust: An R Package for Determining the Relevant Number of Clusters in a Data Set," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 61(i06).
Kohei Adachi & Nickolay T. Trendafilov, 2016. "Sparse principal component analysis subject to prespecified cardinality of loadings," Computational Statistics, Springer, vol. 31(4), pages 1403-1427, December.

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Carlo Cavicchia & Maurizio Vichi & Giorgia Zaccaria, 2023. "Hierarchical disjoint principal component analysis," AStA Advances in Statistical Analysis, Springer;German Statistical Society, vol. 107(3), pages 537-574, September.
Nickolay Trendafilov, 2014. "From simple structure to sparse components: a review," Computational Statistics, Springer, vol. 29(3), pages 431-454, June.
Naoto Yamashita, 2023. "Principal component analysis constrained by layered simple structures," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 17(2), pages 347-367, June.
José Fernando Romero Cañizares & Purificación Vicente Galindo & Yannis Phillis & Evangelos Grigoroudis, 2022. "Graphical sustainability analysis using disjoint biplots," Operational Research, Springer, vol. 22(2), pages 1575-1596, April.
Bolívar, Fernando & Duran, Miguel A. & Lozano-Vivas, Ana, 2023. "Bank business models, size, and profitability," Finance Research Letters, Elsevier, vol. 53(C).
- F. Bolivar & M. A. Duran & A. Lozano-Vivas, 2024. "Bank Business Models, Size, and Profitability," Papers 2401.12323, arXiv.org.
Kim, Hyun Hak & Swanson, Norman R., 2018. "Mining big data using parsimonious factor, machine learning, variable selection and shrinkage methods," International Journal of Forecasting, Elsevier, vol. 34(2), pages 339-354.
Reder, Maik & YÃ¼rÃ¼ÅŸen, Nurseda Y. & Melero, Julio J., 2018. "Data-driven learning framework for associating weather conditions and wind turbine failures," Reliability Engineering and System Safety, Elsevier, vol. 169(C), pages 554-569.
Alfonso Iodice D’Enza & Francesco Palumbo, 2013. "Iterative factor clustering of binary data," Computational Statistics, Springer, vol. 28(2), pages 789-807, April.
Marcin Gąsior, 2021. "Environmental Attitudes and Willingness to Purchase Online—Classification Approach," Sustainability, MDPI, vol. 13(15), pages 1-17, August.
Kohei Adachi & Nickolay T. Trendafilov, 2018. "Sparsest factor analysis for clustering variables: a matrix decomposition approach," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 12(3), pages 559-585, September.
Roopam Shukla & Ankit Agarwal & Kamna Sachdeva & Juergen Kurths & P. K. Joshi, 2019. "Climate change perception: an analysis of climate change and risk perceptions among farmer types of Indian Western Himalayas," Climatic Change, Springer, vol. 152(1), pages 103-119, January.
Yannis Yatracos, 2013. "Detecting Clusters in the Data from Variance Decompositions of Its Projections," Journal of Classification, Springer;The Classification Society, vol. 30(1), pages 30-55, April.
Saemi Shin & Won Suck Yoon & Sang-Hoon Byeon, 2022. "Trends in Occupational Infectious Diseases in South Korea and Classification of Industries According to the Risk of Biological Hazards Using K-Means Clustering," IJERPH, MDPI, vol. 19(19), pages 1-19, September.
Fernández, D. & Arnold, R. & Pledger, S., 2016. "Mixture-based clustering for the ordered stereotype model," Computational Statistics & Data Analysis, Elsevier, vol. 93(C), pages 46-75.
Atsuho Nakayama & Daniel Baier, 2020. "Predicting brand confusion in imagery markets based on deep learning of visual advertisement content," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 14(4), pages 927-945, December.
Blasius, J. & Greenacre, M. & Groenen, P.J.F. & van de Velden, M., 2009. "Special issue on correspondence analysis and related methods," Computational Statistics & Data Analysis, Elsevier, vol. 53(8), pages 3103-3106, June.
Song He & Xinyu Song & Xiaoxi Yang & Jijun Yu & Yuqi Wen & Lianlian Wu & Bowei Yan & Jiannan Feng & Xiaochen Bo, 2021. "COMSUC: A web server for the identification of consensus molecular subtypes of cancer based on multiple methods and multi-omics data," PLOS Computational Biology, Public Library of Science, vol. 17(3), pages 1-10, March.
Jihane El Ouadi & Hanae Errousso & Nicolas Malhene & Siham Benhadou & Hicham Medromi, 2022. "A machine-learning based hybrid algorithm for strategic location of urban bundling hubs to support shared public transport," Quality & Quantity: International Journal of Methodology, Springer, vol. 56(5), pages 3215-3258, October.
Cyril Atkinson-Clement & Eléonore Pigalle, 2021. "What can we learn from Covid-19 pandemic’s impact on human behaviour? The case of France’s lockdown," Palgrave Communications, Palgrave Macmillan, vol. 8(1), pages 1-12, December.
Kreitmair, Ursula & Bower-Bir, Jacob, 2021. "Too different to solve climate change? Experimental evidence on the effects of production and benefit heterogeneity on collective action," Ecological Economics, Elsevier, vol. 184(C).

More about this item

Keywords

Principal component analysis; Clustering of objects; Partitioning of attributes; Semidefinite programming;
All these keywords.

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:stmapp:v:30:y:2021:i:3:d:10.1007_s10260-020-00546-2. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

An empirical comparison of two approaches for CDPCA in high-dimensional data

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data