IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v54y2010i7p1791-1807.html
   My bibliography  Save this article

Prediction of multivariate responses with a selected number of principal components

Author

Listed:
  • Koch, Inge
  • Naito, Kanta

Abstract

This paper proposes a new method and algorithm for predicting multivariate responses in a regression setting. Research into the classification of high dimension low sample size (HDLSS) data, in particular microarray data, has made considerable advances, but regression prediction for high-dimensional data with continuous responses has had less attention. Recently Bair et al. (2006) proposed an efficient prediction method based on supervised principal component regression (PCR). Motivated by the fact that using a larger number of principal components results in better regression performance, this paper extends the method of Bair et al. in several ways: a comprehensive variable ranking is combined with a selection of the best number of components for PCR, and the new method further extends to regression with multivariate responses. The new method is particularly suited to addressing HDLSS problems. Applications to simulated and real data demonstrate the performance of the new method. Comparisons with the findings of Bair et al. (2006) show that for high-dimensional data in particular the new ranking results in a smaller number of predictors and smaller errors.

Suggested Citation

  • Koch, Inge & Naito, Kanta, 2010. "Prediction of multivariate responses with a selected number of principal components," Computational Statistics & Data Analysis, Elsevier, vol. 54(7), pages 1791-1807, July.
  • Handle: RePEc:eee:csdana:v:54:y:2010:i:7:p:1791-1807
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167-9473(10)00045-9
    Download Restriction: Full text for ScienceDirect subscribers only.
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Bair, Eric & Hastie, Trevor & Paul, Debashis & Tibshirani, Robert, 2006. "Prediction by Supervised Principal Components," Journal of the American Statistical Association, American Statistical Association, vol. 101, pages 119-137, March.
    2. Laura J. van 't Veer & Hongyue Dai & Marc J. van de Vijver & Yudong D. He & Augustinus A. M. Hart & Mao Mao & Hans L. Peterse & Karin van der Kooy & Matthew J. Marton & Anke T. Witteveen & George J. S, 2002. "Gene expression profiling predicts clinical outcome of breast cancer," Nature, Nature, vol. 415(6871), pages 530-536, January.
    3. Leo Breiman & Jerome H. Friedman, 1997. "Predicting Multivariate Responses in Multiple Linear Regression," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 59(1), pages 3-54.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Tsay, Ruey S. & Ando, Tomohiro, 2012. "Bayesian panel data analysis for exploring the impact of subprime financial crisis on the US stock market," Computational Statistics & Data Analysis, Elsevier, vol. 56(11), pages 3345-3365.
    2. Tamatani, Mitsuru & Koch, Inge & Naito, Kanta, 2012. "Pattern recognition based on canonical correlations in a high dimension low sample size context," Journal of Multivariate Analysis, Elsevier, vol. 111(C), pages 350-367.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Antoniadis, Anestis & Fryzlewicz, Piotr & Letué, Frédérique, 2010. "The Dantzig selector in Cox's proportional hazards model," LSE Research Online Documents on Economics 30992, London School of Economics and Political Science, LSE Library.
    2. Gaynor, Sheila & Bair, Eric, 2017. "Identification of relevant subtypes via preweighted sparse clustering," Computational Statistics & Data Analysis, Elsevier, vol. 116(C), pages 139-154.
    3. van Wieringen, Wessel N. & Kun, David & Hampel, Regina & Boulesteix, Anne-Laure, 2009. "Survival prediction using gene expression data: A review and comparison," Computational Statistics & Data Analysis, Elsevier, vol. 53(5), pages 1590-1603, March.
    4. Paul Hewson & Keming Yu, 2008. "Quantile regression for binary performance indicators," Applied Stochastic Models in Business and Industry, John Wiley & Sons, vol. 24(5), pages 401-418, September.
    5. Eric Hillebrand & Huiyu Huang & Tae-Hwy Lee & Canlin Li, 2018. "Using the Entire Yield Curve in Forecasting Output and Inflation," Econometrics, MDPI, vol. 6(3), pages 1-27, August.
    6. Tomohiro Ando & Ruey S. Tsay, 2009. "Model selection for generalized linear models with factor‐augmented predictors," Applied Stochastic Models in Business and Industry, John Wiley & Sons, vol. 25(3), pages 207-235, May.
    7. Jewson Stephen & Penzer Jeremy, 2006. "Estimating Trends in Weather Series: Consequences for Pricing Derivatives," Studies in Nonlinear Dynamics & Econometrics, De Gruyter, vol. 10(3), pages 1-17, September.
    8. Luebke, Karsten & Czogiel, Irina & Weihs, Claus, 2004. "Latent Factor Prediction Pursuit for Rank Deficient Regressors," Technical Reports 2004,75, Technische Universität Dortmund, Sonderforschungsbereich 475: Komplexitätsreduktion in multivariaten Datenstrukturen.
    9. Tibshirani Robert J., 2009. "Univariate Shrinkage in the Cox Model for High Dimensional Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 8(1), pages 1-18, April.
    10. Jing Zhang & Qihua Wang & Xuan Wang, 2022. "Surrogate-variable-based model-free feature screening for survival data under the general censoring mechanism," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 74(2), pages 379-397, April.
    11. Kui Shen & Nan Song & Youngchul Kim & Chunqiao Tian & Shara D Rice & Michael J Gabrin & W Fraser Symmans & Lajos Pusztai & Jae K Lee, 2012. "A Systematic Evaluation of Multi-Gene Predictors for the Pathological Response of Breast Cancer Patients to Chemotherapy," PLOS ONE, Public Library of Science, vol. 7(11), pages 1-9, November.
    12. Xiuli Du & Xiaohu Jiang & Jinguan Lin, 2023. "Multinomial Logistic Factor Regression for Multi-source Functional Block-wise Missing Data," Psychometrika, Springer;The Psychometric Society, vol. 88(3), pages 975-1001, September.
    13. Gaorong Li & Liugen Xue & Heng Lian, 2012. "SCAD-penalised generalised additive models with non-polynomial dimensionality," Journal of Nonparametric Statistics, Taylor & Francis Journals, vol. 24(3), pages 681-697.
    14. Zemin Zheng & Jie Zhang & Yang Li, 2022. "L 0 -Regularized Learning for High-Dimensional Additive Hazards Regression," INFORMS Journal on Computing, INFORMS, vol. 34(5), pages 2762-2775, September.
    15. Wang Zhu & Wang C.Y., 2010. "Buckley-James Boosting for Survival Analysis with High-Dimensional Biomarker Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 9(1), pages 1-33, June.
    16. Lian, Heng & Du, Pang & Li, YuanZhang & Liang, Hua, 2014. "Partially linear structure identification in generalized additive models with NP-dimensionality," Computational Statistics & Data Analysis, Elsevier, vol. 80(C), pages 197-208.
    17. Caroline Jardet & Baptiste Meunier, 2022. "Nowcasting world GDP growth with high‐frequency data," Journal of Forecasting, John Wiley & Sons, Ltd., vol. 41(6), pages 1181-1200, September.
    18. Tommaso Proietti, 2016. "On the Selection of Common Factors for Macroeconomic Forecasting," Advances in Econometrics, in: Dynamic Factor Models, volume 35, pages 593-628, Emerald Group Publishing Limited.
    19. Federico Pavone & Juho Piironen & Paul-Christian Bürkner & Aki Vehtari, 2023. "Using reference models in variable selection," Computational Statistics, Springer, vol. 38(1), pages 349-371, March.
    20. Min Cai & Hui Dai & Yongyong Qiu & Yang Zhao & Ruyang Zhang & Minjie Chu & Juncheng Dai & Zhibin Hu & Hongbing Shen & Feng Chen, 2013. "SNP Set Association Analysis for Genome-Wide Association Studies," PLOS ONE, Public Library of Science, vol. 8(5), pages 1-10, May.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:54:y:2010:i:7:p:1791-1807. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.