Estimating residual variance in random forest regression

My bibliography Save this article

Estimating residual variance in random forest regression

Author

Listed:

Mendez, Guillermo
Lohr, Sharon

Registered:

Abstract

Random forest, a data-mining technique which uses multiple classification or regression trees, is a popular algorithm used for prediction. Inference and goodness-of-fit assessment, however, may require an estimator of variability; in many applications the residual variance is of primary interest. This paper proposes two estimators of residual variance for random forest regression that take advantage of byproducts of the algorithm. The first estimator is based on the residual sum of squares from a random forest fit and uses a bootstrap bias correction. The second estimator is a difference-based estimator that uses proximity measures as weights. The estimators are evaluated through Monte Carlo simulations. Applications of the methods to the problem of assessing the relative variability of males and females on cognitive and achievement tests are discussed, and the methods are applied to estimate the residual variance in test scores for male and female students on the mathematics portion of the 2007 Arizona Instrument to Measure Standards.

Suggested Citation

Mendez, Guillermo & Lohr, Sharon, 2011. "Estimating residual variance in random forest regression," Computational Statistics & Data Analysis, Elsevier, vol. 55(11), pages 2937-2950, November.

Handle: RePEc:eee:csdana:v:55:y:2011:i:11:p:2937-2950

Download full text from publisher

As the access to this document is restricted, you may want to search for a different version of it.

References listed on IDEAS

Tiejun Tong & Yuedong Wang, 2005. "Estimating residual variance in nonparametric regression using least squares," Biometrika, Biometrika Trust, vol. 92(4), pages 821-830, December.
Shoemaker L.H., 2003. "Fixing the F Test for Equal Variances," The American Statistician, American Statistical Association, vol. 57, pages 105-114, May.
Bradley Efron, 2004. "The Estimation of Prediction Error: Covariance Penalties and Cross-Validation," Journal of the American Statistical Association, American Statistical Association, vol. 99, pages 619-632, January.
Biau, Gérard & Devroye, Luc, 2010. "On the layered nearest neighbour estimate, the bagged nearest neighbour estimate and the random forest method in regression and classification," Journal of Multivariate Analysis, Elsevier, vol. 101(10), pages 2499-2518, November.
Cahoy, Dexter O., 2010. "A bootstrap test for equality of variances," Computational Statistics & Data Analysis, Elsevier, vol. 54(10), pages 2306-2316, October.
Lin, Yi & Jeon, Yongho, 2006. "Random Forests and Adaptive Nearest Neighbors," Journal of the American Statistical Association, American Statistical Association, vol. 101, pages 578-590, June.
Spokoiny, Vladimir, 2002. "Variance Estimation for High-Dimensional Regression Models," Journal of Multivariate Analysis, Elsevier, vol. 82(1), pages 111-133, July.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

V. Kohestani & M. Hassanlourad & A. Ardakani, 2015. "Evaluation of liquefaction potential based on CPT data using random forest," Natural Hazards: Journal of the International Society for the Prevention and Mitigation of Natural Hazards, Springer;International Society for the Prevention and Mitigation of Natural Hazards, vol. 79(2), pages 1079-1089, November.
Peter Hall & Joel L. Horowitz, 2012. "A simple bootstrap method for constructing nonparametric confidence bands for functions," CeMMAP working papers 14/12, Institute for Fiscal Studies.
Peter Hall & Joel L. Horowitz, 2013. "A simple bootstrap method for constructing nonparametric confidence bands for functions," CeMMAP working papers CWP29/13, Centre for Microdata Methods and Practice, Institute for Fiscal Studies.
Peter Hall & Joel L. Horowitz, 2012. "A simple bootstrap method for constructing nonparametric confidence bands for functions," CeMMAP working papers CWP14/12, Centre for Microdata Methods and Practice, Institute for Fiscal Studies.
Patrick Krennmair & Timo Schmid, 2022. "Flexible domain prediction using mixed effects random forests," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(5), pages 1865-1894, November.
Yihui Chen & Minjie Li, 2019. "Evaluation of influencing factors on tea production based on random forest regression and mean impact value," Agricultural Economics, Czech Academy of Agricultural Sciences, vol. 65(7), pages 340-347.
Ramosaj, Burim & Pauly, Markus, 2019. "Consistent estimation of residual variance with random forest Out-Of-Bag errors," Statistics & Probability Letters, Elsevier, vol. 151(C), pages 49-57.
Peter Hall & Joel L. Horowitz, 2013. "A simple bootstrap method for constructing nonparametric confidence bands for functions," CeMMAP working papers 29/13, Institute for Fiscal Studies.
J Sunil Rao & Erin Kobetz & Huilin Yu & Jordan Baeker-Bispo & Zinzi Bailey, 2023. "Partially Recursively Induced Structured Moderation (PRISM) for modeling racial differences in endometrial cancer survival," PLOS ONE, Public Library of Science, vol. 18(1), pages 1-19, January.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Susan Athey & Julie Tibshirani & Stefan Wager, 2016. "Generalized Random Forests," Papers 1610.01271, arXiv.org, revised Apr 2018.
- Athey, Susan & Tibshirani, Julie & Wager, Stefan, 2017. "Generalized Random Forests," Research Papers 3575, Stanford University, Graduate School of Business.
Jincheng Shen & Lu Wang & Jeremy M. G. Taylor, 2017. "Estimation of the optimal regime in treatment of prostate cancer recurrence from observational data using flexible weighting models," Biometrics, The International Biometric Society, vol. 73(2), pages 635-645, June.
Gérard Biau & Erwan Scornet, 2016. "A random forest guided tour," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 25(2), pages 197-227, June.
Uguccioni, James, 2022. "The long-run effects of parental unemployment in childhood," CLEF Working Paper Series 45, Canadian Labour Economics Forum (CLEF), University of Waterloo.
Guoyi Zhang & Yan Lu, 2012. "Bias-corrected random forests in regression," Journal of Applied Statistics, Taylor & Francis Journals, vol. 39(1), pages 151-160, March.
Liitiäinen, Elia & Corona, Francesco & Lendasse, Amaury, 2010. "Residual variance estimation using a nearest neighbor statistic," Journal of Multivariate Analysis, Elsevier, vol. 101(4), pages 811-823, April.
Ramosaj, Burim & Pauly, Markus, 2019. "Consistent estimation of residual variance with random forest Out-Of-Bag errors," Statistics & Probability Letters, Elsevier, vol. 151(C), pages 49-57.
Wang, WenWu & Yu, Ping, 2017. "Asymptotically optimal differenced estimators of error variance in nonparametric regression," Computational Statistics & Data Analysis, Elsevier, vol. 105(C), pages 125-143.
Zhexiao Lin & Fang Han, 2022. "On regression-adjusted imputation estimators of the average treatment effect," Papers 2212.05424, arXiv.org, revised Jan 2023.
Ramos-Guajardo, Ana Belén & Lubiano, María Asunción, 2012. "K-sample tests for equality of variances of random fuzzy sets," Computational Statistics & Data Analysis, Elsevier, vol. 56(4), pages 956-966.
Theo Dijkstra, 2014. "Ridge regression and its degrees of freedom," Quality & Quantity: International Journal of Methodology, Springer, vol. 48(6), pages 3185-3193, November.
Klugkist, Irene & Hoijtink, Herbert, 2009. "Obtaining similar null distributions in the normal linear model using computational methods," Computational Statistics & Data Analysis, Elsevier, vol. 53(4), pages 877-888, February.
Sieds, 2012. "Complete Volume LXVI n.1 2012," RIEDS - Rivista Italiana di Economia, Demografia e Statistica - The Italian Journal of Economic, Demographic and Statistical Studies, SIEDS Societa' Italiana di Economia Demografia e Statistica, vol. 66(1), pages 1-296.
Jerinsh Jeyapaulraj & Dhruv Desai & Peter Chu & Dhagash Mehta & Stefano Pasquali & Philip Sommer, 2022. "Supervised similarity learning for corporate bonds using Random Forest proximities," Papers 2207.04368, arXiv.org, revised Oct 2022.
Luu, Tung Duy & Fadili, Jalal & Chesneau, Christophe, 2019. "PAC-Bayesian risk bounds for group-analysis sparse regression by exponential weighting," Journal of Multivariate Analysis, Elsevier, vol. 171(C), pages 209-233.
Philip Pallmann & Ludwig Hothorn & Gemechis Djira, 2014. "A Levene-type test of homogeneity of variances against ordered alternatives," Computational Statistics, Springer, vol. 29(6), pages 1593-1608, December.
Jingxin Zhao & Heng Peng & Tao Huang, 2018. "Variance estimation for semiparametric regression models by local averaging," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 27(2), pages 453-476, June.
Joshua Rosaler & Luca Candelori & Vahagn Kirakosyan & Kharen Musaelian & Ryan Samson & Martin T. Wells & Dhagash Mehta & Stefano Pasquali, 2025. "Supervised Similarity for High-Yield Corporate Bonds with Quantum Cognition Machine Learning," Papers 2502.01495, arXiv.org.
Hettihewa, Samanthala & Saha, Shrabani & Zhang, Hanxiong, 2018. "Does an aging population influence stock markets? Evidence from New Zealand," Economic Modelling, Elsevier, vol. 75(C), pages 142-158.
David M. Ritzwoller & Vasilis Syrgkanis, 2024. "Simultaneous Inference for Local Structural Parameters with Random Forests," Papers 2405.07860, arXiv.org, revised Sep 2024.

More about this item

Keywords

Bootstrap Gender gap Greater male variability hypothesis Nonparametric regression Proximity measure Regression tree Sex differences;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:55:y:2011:i:11:p:2937-2950. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Estimating residual variance in random forest regression

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data