Bollinger, Christopher R. (University of Kentucky) Chandra, Amitabh () (Dartmouth College, NBER and IZA Bonn)
Abstract
In empirical research it is common practice to use sensible rules of thumb for cleaning data. Measurement error is often the justification for removing (trimming) or recoding (winsorizing) observations whose values lie outside a specified range. We consider a general measurement error process that nests many plausible models. Analytic results demonstrate that winsorizing and trimming are only solutions for a narrow class of measurement error processes. Indeed, for the measurement error processes found in most social-science data, such procedures can induce or exacerbate bias, and even inflate the variance estimates. We term this source of bias “Iatrogenic” (or econometrician induced) error. Monte Carlo simulations and empirical results from the Census PUMS data and 2001 CPS data demonstrate the fragility of trimming and winsorizing as solutions to measurement error in the dependent variable. Even on asymptotic variance and RMSE criteria, we are unable to find generalizable justifications for commonly used cleaning procedures.
Download Info
To download:
If you experience problems downloading a file, check if you have the
proper application to
view it first. Information about this may be contained
in the File-Format links below. In case of further problems read
the IDEAS help
page. Note that these files are not on the IDEAS
site. Please be patient as the files may be large.
Publisher Info
Paper provided by Institute for the Study of Labor (IZA) in its series IZA Discussion Papers with number
1093.
Find related papers by JEL classification: C1 - Mathematical and Quantitative Methods - - Econometric and Statistical Methods: General J1 - Labor and Demographic Economics - - Demographic Economics
This paper has been announced in the following NEP Reports:
References listed on IDEAS Please report citation or reference errors to , or , if you are the registered author of the cited work, log in to your RePEc Author Service profile, click on "citations" and make appropriate adjustments.: