IDEAS home Printed from https://ideas.repec.org/a/bpj/sagmbi/v7y2008i1n12.html

Adapting Prediction Error Estimates for Biased Complexity Selection in High-Dimensional Bootstrap Samples

Author

Listed:
  • Binder Harald

    (Institute of Medical Biometry and Medical Informatics, University Medical Center Freiburg)

  • Schumacher Martin

    (Institute of Medical Biometry and Medical Informatics, University Medical Center Freiburg)

Abstract

The bootstrap is a tool that allows for efficient evaluation of prediction performance of statistical techniques without having to set aside data for validation. This is especially important for high-dimensional data, e.g., arising from microarrays, because there the number of observations is often limited. For avoiding overoptimism the statistical technique to be evaluated has to be applied to every bootstrap sample in the same manner it would be used on new data. This includes a selection of complexity, e.g., the number of boosting steps for gradient boosting algorithms. Using the latter, we demonstrate in a simulation study that complexity selection in conventional bootstrap samples, drawn with replacement, is severely biased in many scenarios. This translates into a considerable bias of prediction error estimates, often underestimating the amount of information that can be extracted from high-dimensional data. Potential remedies for this complexity selection bias, such as alternatively using a fixed level of complexity or of using sampling without replacement are investigated and it is shown that the latter works well in many settings. We focus on high-dimensional binary response data, with bootstrap .632+ estimates of the Brier score for performance evaluation, and censored time-to-event data with .632+ prediction error curve estimates. The latter, with the modified bootstrap procedure, is then applied to an example with microarray data from patients with diffuse large B-cell lymphoma.

Suggested Citation

  • Binder Harald & Schumacher Martin, 2008. "Adapting Prediction Error Estimates for Biased Complexity Selection in High-Dimensional Bootstrap Samples," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 7(1), pages 1-28, March.
  • Handle: RePEc:bpj:sagmbi:v:7:y:2008:i:1:n:12
    DOI: 10.2202/1544-6115.1346
    as

    Download full text from publisher

    File URL: https://doi.org/10.2202/1544-6115.1346
    Download Restriction: For access to full text, subscription to the journal or payment for the individual article is required.

    File URL: https://libkey.io/10.2202/1544-6115.1346?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to

    for a different version of it.

    References listed on IDEAS

    as
    1. Thomas A. Gerds & Martin Schumacher, 2007. "Efron-Type Measures of Prediction Error for Survival Analysis," Biometrics, The International Biometric Society, vol. 63(4), pages 1283-1287, December.
    2. Buhlmann P. & Yu B., 2003. "Boosting With the L2 Loss: Regression and Classification," Journal of the American Statistical Association, American Statistical Association, vol. 98, pages 324-339, January.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Christine Porzelius & Martin Schumacher & Harald Binder, 2011. "The benefit of data-based model complexity selection via prediction error curves in time-to-event data," Computational Statistics, Springer, vol. 26(2), pages 293-302, June.
    2. Mogensen, Ulla B. & Ishwaran, Hemant & Gerds, Thomas A., 2012. "Evaluating Random Forests for Survival Analysis Using Prediction Error Curves," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 50(i11).
    3. Lore Zumeta-Olaskoaga & Maximilian Weigert & Jon Larruskain & Eder Bikandi & Igor Setuain & Josean Lekue & Helmut Küchenhoff & Dae-Jin Lee, 2023. "Prediction of sports injuries in football: a recurrent time-to-event approach using regularized Cox models," AStA Advances in Statistical Analysis, Springer;German Statistical Society, vol. 107(1), pages 101-126, March.
    4. Sill, Martin & Hielscher, Thomas & Becker, Natalia & Zucknick, Manuela, 2014. "c060: Extended Inference with Lasso and Elastic-Net Regularized Cox and Generalized Linear Models," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 62(i05).
    5. Bernd Bischl & Julia Schiffner & Claus Weihs, 2013. "Benchmarking local classification methods," Computational Statistics, Springer, vol. 28(6), pages 2599-2619, December.
    6. Stefanie Hieke & Axel Benner & Richard F Schlenk & Martin Schumacher & Lars Bullinger & Harald Binder, 2016. "Identifying Prognostic SNPs in Clinical Cohorts: Complementing Univariate Analyses by Resampling and Multivariable Modeling," PLOS ONE, Public Library of Science, vol. 11(5), pages 1-18, May.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Tutz, Gerhard & Pößnecker, Wolfgang & Uhlmann, Lorenz, 2015. "Variable selection in general multinomial logit models," Computational Statistics & Data Analysis, Elsevier, vol. 82(C), pages 207-222.
    2. Ivan Chang, Yuan-Chin & Huang, Yufen & Huang, Yu-Pai, 2010. "Early stopping in L2Boosting," Computational Statistics & Data Analysis, Elsevier, vol. 54(10), pages 2203-2213, October.
    3. Gerhard Tutz & Moritz Berger, 2018. "Tree-structured modelling of categorical predictors in generalized additive regression," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 12(3), pages 737-758, September.
    4. Bissantz, Nicolai & Hohage, T. & Munk, Axel & Ruymgaart, F., 2007. "Convergence rates of general regularization methods for statistical inverse problems and applications," Technical Reports 2007,04, Technische Universität Dortmund, Sonderforschungsbereich 475: Komplexitätsreduktion in multivariaten Datenstrukturen.
    5. Kea BARET, 2021. "Fiscal rules’ compliance and Social Welfare," Working Papers of BETA 2021-38, Bureau d'Economie Théorique et Appliquée, UDS, Strasbourg.
    6. Mittnik, Stefan & Robinzonov, Nikolay & Spindler, Martin, 2015. "Stock market volatility: Identifying major drivers and the nature of their impact," Journal of Banking & Finance, Elsevier, vol. 58(C), pages 1-14.
    7. Shafik, Nivien & Tutz, Gerhard, 2009. "Boosting nonlinear additive autoregressive time series," Computational Statistics & Data Analysis, Elsevier, vol. 53(7), pages 2453-2464, May.
    8. Wang Zhu & Wang C.Y., 2010. "Buckley-James Boosting for Survival Analysis with High-Dimensional Biomarker Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 9(1), pages 1-33, June.
    9. Kevin He & Ji Zhu & Jian Kang & Yi Li, 2022. "Stratified Cox models with time‐varying effects for national kidney transplant patients: A new blockwise steepest ascent method," Biometrics, The International Biometric Society, vol. 78(3), pages 1221-1232, September.
    10. Tutz, Gerhard & Leitenstorfer, Florian, 2006. "Response shrinkage estimators in binary regression," Computational Statistics & Data Analysis, Elsevier, vol. 50(10), pages 2878-2901, June.
    11. Leitenstorfer, Florian & Tutz, Gerhard, 2007. "Knot selection by boosting techniques," Computational Statistics & Data Analysis, Elsevier, vol. 51(9), pages 4605-4621, May.
    12. Martijn Kagie & Michiel Van Wezel, 2007. "Hedonic price models and indices based on boosting applied to the Dutch housing market," Intelligent Systems in Accounting, Finance and Management, John Wiley & Sons, Ltd., vol. 15(3‐4), pages 85-106, July.
    13. Lin, Yi, 2004. "A note on margin-based loss functions in classification," Statistics & Probability Letters, Elsevier, vol. 68(1), pages 73-82, June.
    14. Sigrist, Fabio & Hirnschall, Christoph, 2019. "Grabit: Gradient tree-boosted Tobit models for default prediction," Journal of Banking & Finance, Elsevier, vol. 102(C), pages 177-192.
    15. Ng, Serena, 2013. "Variable Selection in Predictive Regressions," Handbook of Economic Forecasting, in: G. Elliott & C. Granger & A. Timmermann (ed.), Handbook of Economic Forecasting, edition 1, volume 2, chapter 0, pages 752-789, Elsevier.
    16. Hofner, Benjamin & Mayr, Andreas & Schmid, Matthias, 2016. "gamboostLSS: An R Package for Model Building and Variable Selection in the GAMLSS Framework," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 74(i01).
    17. Marra, Giampiero & Wood, Simon N., 2011. "Practical variable selection for generalized additive models," Computational Statistics & Data Analysis, Elsevier, vol. 55(7), pages 2372-2387, July.
    18. Sariyar Murat & Schumacher Martin & Binder Harald, 2014. "A boosting approach for adapting the sparsity of risk prediction signatures based on different molecular levels," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 13(3), pages 343-357, June.
    19. Robin Van Oirbeek & Emmanuel Lesaffre, 2018. "An Investigation of the Discriminatory Ability of the Clustering Effect of the Frailty Survival Model," Biostatistics and Biometrics Open Access Journal, Juniper Publishers Inc., vol. 6(3), pages 87-98, April.
    20. Ziwei Mei & Zhentao Shi & Peter C. B. Phillips, 2022. "The boosted HP filter is more general than you might think," Cowles Foundation Discussion Papers 2348, Cowles Foundation for Research in Economics, Yale University.

    More about this item

    Keywords

    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bpj:sagmbi:v:7:y:2008:i:1:n:12. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Peter Golla (email available below). General contact details of provider: https://www.degruyterbrill.com .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.