IDEAS home Printed from https://ideas.repec.org/a/spr/advdac/v12y2018i2d10.1007_s11634-017-0296-8.html
   My bibliography  Save this article

Clusterwise analysis for multiblock component methods

Author

Listed:
  • Stéphanie Bougeard

    (Anses (French agency for food, environmental and occupational health safety))

  • Hervé Abdi

    (The University of Texas at Dallas)

  • Gilbert Saporta

    (CEDRIC CNAM)

  • Ndèye Niang

    (CEDRIC CNAM)

Abstract

Multiblock component methods are applied to data sets for which several blocks of variables are measured on a same set of observations with the goal to analyze the relationships between these blocks of variables. In this article, we focus on multiblock component methods that integrate the information found in several blocks of explanatory variables in order to describe and explain one set of dependent variables. In the following, multiblock PLS and multiblock redundancy analysis are chosen, as particular cases of multiblock component methods when one set of variables is explained by a set of predictor variables that is organized into blocks. Because these multiblock techniques assume that the observations come from a homogeneous population they will provide suboptimal results when the observations actually come from different populations. A strategy to palliate this problem—presented in this article—is to use a technique such as clusterwise regression in order to identify homogeneous clusters of observations. This approach creates two new methods that provide clusters that have their own sets of regression coefficients. This combination of clustering and regression improves the overall quality of the prediction and facilitates the interpretation. In addition, the minimization of a well-defined criterion—by means of a sequential algorithm—ensures that the algorithm converges monotonously. Finally, the proposed method is distribution-free and can be used when the explanatory variables outnumber the observations within clusters. The proposed clusterwise multiblock methods are illustrated with of a simulation study and a (simulated) example from marketing.

Suggested Citation

  • Stéphanie Bougeard & Hervé Abdi & Gilbert Saporta & Ndèye Niang, 2018. "Clusterwise analysis for multiblock component methods," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 12(2), pages 285-313, June.
  • Handle: RePEc:spr:advdac:v:12:y:2018:i:2:d:10.1007_s11634-017-0296-8
    DOI: 10.1007/s11634-017-0296-8
    as

    Download full text from publisher

    File URL: http://link.springer.com/10.1007/s11634-017-0296-8
    File Function: Abstract
    Download Restriction: Access to the full text of the articles in this series is restricted.

    File URL: https://libkey.io/10.1007/s11634-017-0296-8?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Donatella Vicari & Maurizio Vichi, 2013. "Multivariate linear regression for heterogeneous data," Journal of Applied Statistics, Taylor & Francis Journals, vol. 40(6), pages 1209-1230, June.
    2. Esposito Vinzi, Vincenzo & Ringle, Christian M. & Squillacciotti, Silvia & Trinchera, Laura, 2007. "Capturing and Treating Unobserved Heterogeneity by Response Based Segmentation in PLS Path Modeling. A Comparison of Alternative Methods by Computational Experiments," ESSEC Working Papers DR 07019, ESSEC Research Center, ESSEC Business School.
    3. Heungsun Hwang & Wayne Desarbo & Yoshio Takane, 2007. "Fuzzy Clusterwise Generalized Structured Component Analysis," Psychometrika, Springer;The Psychometric Society, vol. 72(2), pages 181-198, June.
    4. Michel Tenenhaus & Arthur Tenenhaus, 2011. "Regularized Generalized Canonical Correlation Analysis," Post-Print hal-00609220, HAL.
    5. Preda, C. & Saporta, G., 2005. "Clusterwise PLS regression on a stochastic process," Computational Statistics & Data Analysis, Elsevier, vol. 49(1), pages 99-108, April.
    6. Carsten Hahn & Michael D. Johnson & Andreas Herrmann & Frank Huber, 2002. "Capturing Customer Heterogeneity Using A Finite Mixture Pls Approach," Schmalenbach Business Review (sbr), LMU Munich School of Management, vol. 54(3), pages 243-269, July.
    7. Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
    8. Wayne DeSarbo & William Cron, 1988. "A maximum likelihood methodology for clusterwise linear regression," Journal of Classification, Springer;The Classification Society, vol. 5(2), pages 249-282, September.
    9. Schlittgen, Rainer & Ringle, Christian M. & Sarstedt, Marko & Becker, Jan-Michael, 2016. "Segmentation of PLS path models by iterative reweighted regressions," Journal of Business Research, Elsevier, vol. 69(10), pages 4583-4592.
    10. Preda, C. & Saporta, G., 2005. "PLS regression on a stochastic process," Computational Statistics & Data Analysis, Elsevier, vol. 48(1), pages 149-158, January.
    11. Michel Tenenhaus, 2011. "Regularized generalized canonical correlation analysis," Post-Print hal-00578321, HAL.
    12. Heungsun Hwang & Yoshio Takane, 2004. "Generalized structured component analysis," Psychometrika, Springer;The Psychometric Society, vol. 69(1), pages 81-99, March.
    13. Arthur Tenenhaus & Michel Tenenhaus, 2011. "Regularized Generalized Canonical Correlation Analysis," Psychometrika, Springer;The Psychometric Society, vol. 76(2), pages 257-284, April.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Fei Liu & L. Billard, 2022. "Partition of Interval-Valued Observations Using Regression," Journal of Classification, Springer;The Classification Society, vol. 39(1), pages 55-77, March.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Xavier Bry & Ndèye Niang & Thomas Verron & Stéphanie Bougeard, 2023. "Clusterwise elastic-net regression based on a combined information criterion," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 17(1), pages 75-107, March.
    2. Heungsun Hwang & Gyeongcheol Cho, 2020. "Global Least Squares Path Modeling: A Full-Information Alternative to Partial Least Squares Path Modeling," Psychometrika, Springer;The Psychometric Society, vol. 85(4), pages 947-972, December.
    3. Joki, Kaisa & Bagirov, Adil M. & Karmitsa, Napsu & Mäkelä, Marko M. & Taheri, Sona, 2020. "Clusterwise support vector linear regression," European Journal of Operational Research, Elsevier, vol. 287(1), pages 19-35.
    4. Husson, François & Josse, Julie & Saporta, Gilbert, 2016. "Jan de Leeuw and the French School of Data Analysis," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 73(i06).
    5. Xiuli Du & Xiaohu Jiang & Jinguan Lin, 2023. "Multinomial Logistic Factor Regression for Multi-source Functional Block-wise Missing Data," Psychometrika, Springer;The Psychometric Society, vol. 88(3), pages 975-1001, September.
    6. Wang, Wenjia & Zhou, Yi-Hui, 2021. "Eigenvector-based sparse canonical correlation analysis: Fast computation for estimation of multiple canonical vectors," Journal of Multivariate Analysis, Elsevier, vol. 185(C).
    7. Evermann, Joerg & Tate, Mary, 2016. "Assessing the predictive performance of structural equation model estimators," Journal of Business Research, Elsevier, vol. 69(10), pages 4565-4582.
    8. Adil M. Bagirov & Julien Ugon & Hijran G. Mirzayeva, 2015. "Nonsmooth Optimization Algorithm for Solving Clusterwise Linear Regression Problems," Journal of Optimization Theory and Applications, Springer, vol. 164(3), pages 755-780, March.
    9. Sarstedt, Marko & Hair, Joseph F. & Ringle, Christian M. & Thiele, Kai O. & Gudergan, Siegfried P., 2016. "Estimation issues with PLS and CBSEM: Where the bias lies!," Journal of Business Research, Elsevier, vol. 69(10), pages 3998-4010.
    10. Tenenhaus, Arthur & Philippe, Cathy & Frouin, Vincent, 2015. "Kernel Generalized Canonical Correlation Analysis," Computational Statistics & Data Analysis, Elsevier, vol. 90(C), pages 114-131.
    11. Michel Tenenhaus & Arthur Tenenhaus & Patrick J. F. Groenen, 2017. "Regularized Generalized Canonical Correlation Analysis: A Framework for Sequential Multiblock Component Methods," Psychometrika, Springer;The Psychometric Society, vol. 82(3), pages 737-777, September.
    12. Joseph F. Hair & G. Tomas M. Hult & Christian M. Ringle & Marko Sarstedt & Kai Oliver Thiele, 2017. "Mirror, mirror on the wall: a comparative evaluation of composite-based structural equation modeling methods," Journal of the Academy of Marketing Science, Springer, vol. 45(5), pages 616-632, September.
    13. Rosaria Romano & Francesco Palumbo, 2021. "Partial possibilistic regression path modeling: handling uncertainty in path modeling," Computational Statistics, Springer, vol. 36(1), pages 615-639, March.
    14. Florian Rohart & Benoît Gautier & Amrit Singh & Kim-Anh Lê Cao, 2017. "mixOmics: An R package for ‘omics feature selection and multiple data integration," PLOS Computational Biology, Public Library of Science, vol. 13(11), pages 1-19, November.
    15. Lukáš Malec & Vladimír Janovský, 2020. "Connecting the multivariate partial least squares with canonical analysis: a path-following approach," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 14(3), pages 589-609, September.
    16. Shen, Cencheng & Sun, Ming & Tang, Minh & Priebe, Carey E., 2014. "Generalized canonical correlation analysis for classification," Journal of Multivariate Analysis, Elsevier, vol. 130(C), pages 310-322.
    17. Jan-Michael Becker & Christian Ringle & Marko Sarstedt & Franziska Völckner, 2015. "How collinearity affects mixture regression results," Marketing Letters, Springer, vol. 26(4), pages 643-659, December.
    18. Cristina Davino & Vincenzo Esposito Vinzi, 2016. "Quantile composite-based path modeling," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 10(4), pages 491-520, December.
    19. Tom Frans Wilderjans & Eva Gaer & Henk A. L. Kiers & Iven Mechelen & Eva Ceulemans, 2017. "Principal Covariates Clusterwise Regression (PCCR): Accounting for Multicollinearity and Population Heterogeneity in Hierarchically Organized Data," Psychometrika, Springer;The Psychometric Society, vol. 82(1), pages 86-111, March.
    20. Cristina Davino & Pasquale Dolce & Stefania Taralli & Domenico Vistocco, 2022. "Composite-Based Path Modeling for Conditional Quantiles Prediction. An Application to Assess Health Differences at Local Level in a Well-Being Perspective," Social Indicators Research: An International and Interdisciplinary Journal for Quality-of-Life Measurement, Springer, vol. 161(2), pages 907-936, June.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:advdac:v:12:y:2018:i:2:d:10.1007_s11634-017-0296-8. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.