A nonparametric test for equality of distributions with mixed categorical and continuous data
In this paper we consider the problem of testing for equality of two density or two conditional density functions defined over mixed discrete and continuous variables. We smooth both the discrete and continuous variables, with the smoothing parameters chosen via least-squares cross-validation. The test statistics are shown to have (asymptotic) normal null distributions. However, we advocate the use of bootstrap methods in order to better approximate their null distribution in finite-sample settings and we provide asymptotic validity of the proposed bootstrap method. Simulations show that the proposed tests have better power than both conventional frequency-based tests and smoothing tests based on ad hoc smoothing parameter selection, while a demonstrative empirical application to the joint distribution of earnings and educational attainment underscores the utility of the proposed approach in mixed data settings.
If you experience problems downloading a file, check if you have the proper application to view it first. In case of further problems read the IDEAS help page. Note that these files are not on the IDEAS site. Please be patient as the files may be large.
As the access to this document is restricted, you may want to look for a different version under "Related research" (further below) or search for a different version of it.
References listed on IDEAS
Please report citation or reference errors to , or , if you are the registered author of the cited work, log in to your RePEc Author Service profile, click on "citations" and make appropriate adjustments.:
- Kiefer, Nicholas M. & Racine, Jeffrey S., 2008. "The Smooth Colonel Meets the Reverend," Working Papers 08-01, Cornell University, Center for Analytic Economics.
- Peter Hall & Qi Li & Jeffrey S. Racine, 2007. "Nonparametric Estimation of Regression Functions in the Presence of Irrelevant Regressors," The Review of Economics and Statistics, MIT Press, vol. 89(4), pages 784-789, November.
- Gordon Anderson, 2001. "The Power And Size Of Nonparametric Tests For Common Distributional Characteristics," Econometric Reviews, Taylor & Francis Journals, vol. 20(1), pages 1-30.
- Fan, Yanqin & Li, Qi, 2000. "Consistent Model Specification Tests," Econometric Theory, Cambridge University Press, vol. 16(06), pages 1016-1041, December.
- Ahmad, Ibrahim A. & Li, Qi, 1997. "Testing independence by nonparametric kernel method," Statistics & Probability Letters, Elsevier, vol. 34(2), pages 201-210, June.
- Racine, Jeffrey S. & Maasoumi, Esfandiar, 2007. "A versatile and robust metric entropy test of time-reversibility, and other hypotheses," Journal of Econometrics, Elsevier, vol. 138(2), pages 547-567, June.
- Russell Davidson & James G. MacKinnon, 2001.
"Bootstrap Tests: How Many Bootstraps?,"
1036, Queen's University, Department of Economics.
- Fan, Yanqin, 1998. "Goodness-Of-Fit Tests Based On Kernel Density Estimators With Fixed Smoothing Parameters," Econometric Theory, Cambridge University Press, vol. 14(05), pages 604-621, October.
- Hall, Peter, 1984. "Central limit theorem for integrated square error of multivariate nonparametric density estimators," Journal of Multivariate Analysis, Elsevier, vol. 14(1), pages 1-16, February.
- Robinson, P M, 1991. "Consistent Nonparametric Entropy-Based Testing," Review of Economic Studies, Wiley Blackwell, vol. 58(3), pages 437-53, May.
- Peter Hall & Jeff Racine & Qi Li, 2004. "Cross-Validation and the Estimation of Conditional Probability Densities," Journal of the American Statistical Association, American Statistical Association, vol. 99, pages 1015-1026, December.
- Anderson, N. H. & Hall, P. & Titterington, D. M., 1994. "Two-Sample Test Statistics for Measuring Discrepancies Between Two Multivariate Probability Density Functions Using Kernel-Based Density Estimates," Journal of Multivariate Analysis, Elsevier, vol. 50(1), pages 41-54, July.
- Yongmiao Hong & Halbert White, 2005. "Asymptotic Distribution Theory for Nonparametric Entropy Measures of Serial Dependence," Econometrica, Econometric Society, vol. 73(3), pages 837-901, 05.
- Grund, B. & Hall, P., 1993. "On the Performance of Kernel Estimators for High-Dimensional, Sparse Binary Data," Journal of Multivariate Analysis, Elsevier, vol. 44(2), pages 321-344, February.
- Li, Qi & Racine, Jeff, 2003. "Nonparametric estimation of distributions with categorical and continuous data," Journal of Multivariate Analysis, Elsevier, vol. 86(2), pages 266-292, August.
When requesting a correction, please mention this item's handle: RePEc:eee:econom:v:148:y:2009:i:2:p:186-200. See general information about how to correct material in RePEc.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: (Zhang, Lei)
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
If references are entirely missing, you can add them using this form.
If the full references list an item that is present in RePEc, but the system did not link to it, you can help with this form.
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your profile, as there may be some citations waiting for confirmation.
Please note that corrections may take a couple of weeks to filter through the various RePEc services.