IDEAS home Printed from https://ideas.repec.org/p/tse/wpaper/28126.html
   My bibliography  Save this paper

Stable variable selection for right censored data: comparison of methods

Author

Listed:
  • Besse, Philippe
  • Leconte, Eve
  • Walschaerts, Marie

Abstract

The instability in the selection of models is a major concern with data sets containing a large number of covariates. This paper deals with variable selection methodology in the case of high-dimensional problems where the response variable can be right censored. We focuse on new stable variable selection methods based on bootstrap for two different methodologies commonly used in survival analysis: the Cox proportional hazard model and survival trees. As far as the Cox model is concerned, we investigate the bootstrapping applied to two variable selection techniques: the stepwise algorithm based on the AIC criterion and the L1-penalization of Lasso. Regarding survival trees, we review two methodologies: the bootstrap node-level stabilization and random survival forests. We apply these different approaches to two real data sets, a classical breast cancer data set and an original infertility data set. We compare the methods on two criteria: the prediction error rate based on the Harrell concordance index and the relevance of the interpretation of the corresponding selected models, focusing on the original infertility data set. The aim is to find a compromise between a good prediction performance and ease to interpretation for clinicians. Results suggest that in the case of a small number of individuals, a bootstrapping adapted to L1-penalization in the Cox model or a bootstrap node-level stabilization in survival trees give a good alternative to the random survival forest methodology, known to give the smallest prediction error rate but difficult to interprete by non-statisticians. In a clinical perspective, the complementarity between the methods based on the Cox model and those based on survival trees would permit to built reliable models easy to interprete by the clinician.

Suggested Citation

  • Besse, Philippe & Leconte, Eve & Walschaerts, Marie, 2012. "Stable variable selection for right censored data: comparison of methods," TSE Working Papers 12-486, Toulouse School of Economics (TSE).
  • Handle: RePEc:tse:wpaper:28126
    as

    Download full text from publisher

    File URL: http://arxiv.org/pdf/1203.4928v1.pdf
    File Function: Full text
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Nicolai Meinshausen & Peter Bühlmann, 2010. "Stability selection," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 72(4), pages 417-473, September.
    2. Ciampi, Antonio & Thiffault, Johanne & Nakache, Jean-Pierre & Asselain, Bernard, 1986. "Stratification by stepwise regression, correspondence analysis and recursive partition: a comparison of three methods of analysis for survival data with covariates," Computational Statistics & Data Analysis, Elsevier, vol. 4(3), pages 185-204, October.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Mozhgan Safe & Hossein Mahjub & Javad Faradmal, 2017. "A Comparative Study for Modelling the Survival of Breast Cancer Patients in the West of Iran," Global Journal of Health Science, Canadian Center of Science and Education, vol. 9(2), pages 215-215, February.
    2. Khan Md Hasinur Rahaman & Bhadra Anamika & Howlader Tamanna, 2019. "Stability selection for lasso, ridge and elastic net implemented with AFT models," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 18(5), pages 1-14, October.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Wei, Jie & Chen, Hui, 2020. "Determining the number of factors in approximate factor models by twice K-fold cross validation," Economics Letters, Elsevier, vol. 191(C).
    2. Xiaogang Su & Juanjuan Fan, 2004. "Multivariate Survival Trees: A Maximum Likelihood Approach Based on Frailty Models," Biometrics, The International Biometric Society, vol. 60(1), pages 93-99, March.
    3. Capanu, Marinela & Giurcanu, Mihai & Begg, Colin B. & Gönen, Mithat, 2023. "Subsampling based variable selection for generalized linear models," Computational Statistics & Data Analysis, Elsevier, vol. 184(C).
    4. Gautier Marti & Frank Nielsen & Philippe Donnat & S'ebastien Andler, 2016. "On clustering financial time series: a need for distances between dependent random variables," Papers 1603.07822, arXiv.org.
    5. Skripnikov, A. & Michailidis, G., 2019. "Joint estimation of multiple network Granger causal models," Econometrics and Statistics, Elsevier, vol. 10(C), pages 120-133.
    6. Tan, Xin Lu, 2019. "Optimal estimation of slope vector in high-dimensional linear transformation models," Journal of Multivariate Analysis, Elsevier, vol. 169(C), pages 179-204.
    7. Un Jung Lee & ShengLi Tzeng & Yu-Chuan Chen & James J Chen, 2017. "Development of Predictive Signatures for Treatment Selection in Precision Medicine," Biostatistics and Biometrics Open Access Journal, Juniper Publishers Inc., vol. 2(4), pages 83-88, August.
    8. Yan Zhou & John McArdle, 2015. "Rationale and Applications of Survival Tree and Survival Ensemble Methods," Psychometrika, Springer;The Psychometric Society, vol. 80(3), pages 811-833, September.
    9. Yu, Dengdeng & Zhang, Li & Mizera, Ivan & Jiang, Bei & Kong, Linglong, 2019. "Sparse wavelet estimation in quantile regression with multiple functional predictors," Computational Statistics & Data Analysis, Elsevier, vol. 136(C), pages 12-29.
    10. Sohrabi, Narges & Movaghari, Hadi, 2020. "Reliable factors of Capital structure: Stability selection approach," The Quarterly Review of Economics and Finance, Elsevier, vol. 77(C), pages 296-310.
    11. Soyeon Kim & Veerabhadran Baladandayuthapani & J. Jack Lee, 2017. "Prediction-Oriented Marker Selection (PROMISE): With Application to High-Dimensional Regression," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 9(1), pages 217-245, June.
    12. Jin, Sainan & Miao, Ke & Su, Liangjun, 2021. "On factor models with random missing: EM estimation, inference, and cross validation," Journal of Econometrics, Elsevier, vol. 222(1), pages 745-777.
    13. Susan Athey & Julie Tibshirani & Stefan Wager, 2016. "Generalized Random Forests," Papers 1610.01271, arXiv.org, revised Apr 2018.
    14. Hsu, David, 2015. "Identifying key variables and interactions in statistical models of building energy consumption using regularization," Energy, Elsevier, vol. 83(C), pages 144-155.
    15. Chun, Hyonho & Lee, Myung Hee & Fleet, James C. & Oh, Ji Hwan, 2016. "Graphical models via joint quantile regression with component selection," Journal of Multivariate Analysis, Elsevier, vol. 152(C), pages 162-171.
    16. Guo, Peiyang & Lam, Jacqueline C.K. & Li, Victor O.K., 2019. "Drivers of domestic electricity users’ price responsiveness: A novel machine learning approach," Applied Energy, Elsevier, vol. 235(C), pages 900-913.
    17. De Bin, Riccardo & Boulesteix, Anne-Laure & Sauerbrei, Willi, 2017. "Detection of influential points as a byproduct of resampling-based variable selection procedures," Computational Statistics & Data Analysis, Elsevier, vol. 116(C), pages 19-31.
    18. Solari, Aldo & Djordjilović, Vera, 2022. "Multi split conformal prediction," Statistics & Probability Letters, Elsevier, vol. 184(C).
    19. Jayachandran, Seema & Biradavolu, Monica & Cooper, Jan, 2023. "Using machine learning and qualitative interviews to design a five-question survey module for women’s agency," World Development, Elsevier, vol. 161(C).
    20. Raheem, S.M. Enayetur & Ahmed, S. Ejaz & Doksum, Kjell A., 2012. "Absolute penalty and shrinkage estimation in partially linear models," Computational Statistics & Data Analysis, Elsevier, vol. 56(4), pages 874-891.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:tse:wpaper:28126. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: the person in charge (email available below). General contact details of provider: https://edirc.repec.org/data/tsetofr.html .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.