IDEAS home Printed from https://ideas.repec.org/a/csb/stintr/v17y2016i3p429-447.html
   My bibliography  Save this article

Prediction of a Function of Misclassified Binary Data

Author

Listed:
  • Partha Lahiri
  • Noriah M. Al-Kandari

Abstract

We consider the problem of predicting a function of misclassified binary variables. We make an interesting observation that the naive predictor, which ignores the misclassification errors, is unbiased even if the total misclassification error is high as long as the probabilities of false positives and false negatives are identical. Other than this case, the bias of the naive predictor depends on the misclassification distribution and the magnitude of the bias can be high in certain cases. We correct the bias of the naive predictor using a double sampling idea where both inaccurate and accurate measurements are taken on the binary variable for all the units of a sample drawn from the original data using a probability sampling scheme. Using this additional information and design-based sample survey theory, we derive a biascorrected predictor. We examine the cases where the new bias-corrected predictors can also improve over the naive predictor in terms of mean square error (MSE).

Suggested Citation

  • Partha Lahiri & Noriah M. Al-Kandari, 2016. "Prediction of a Function of Misclassified Binary Data," Statistics in Transition new series, Główny Urząd Statystyczny (Polska), vol. 17(3), pages 429-447, September.
  • Handle: RePEc:csb:stintr:v:17:y:2016:i:3:p:429-447
    as

    Download full text from publisher

    File URL: http://index.stat.gov.pl/repec/files/csb/stintr/csb_stintr_v17_2016_i3_n5.pdf
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. P. Lahiri & Michael D. Larsen, 2005. "Regression Analysis With Linked Data," Journal of the American Statistical Association, American Statistical Association, vol. 100, pages 222-230, March.
    2. Dewi Rahardja & Ying Yang, 2015. "Maximum likelihood estimation of a binomial proportion using one-sample misclassified binary data," Statistica Neerlandica, Netherlands Society for Statistics and Operations Research, vol. 69(3), pages 272-280, August.
    3. Paul Gustafson & Nhu D. Le & Refik Saskin, 2001. "Case–Control Analysis with Partial Knowledge of Exposure Misclassification Probabilities," Biometrics, The International Biometric Society, vol. 57(2), pages 598-609, June.
    4. Boese, Doyle H. & Young, Dean M. & Stamey, James D., 2006. "Confidence intervals for a binomial parameter based on binary data subject to false-positive misclassification," Computational Statistics & Data Analysis, Elsevier, vol. 50(12), pages 3369-3385, August.
    5. Anil Gaba & Robert L. Winkler, 1992. "Implications of Errors in Survey Data: A Bayesian Model," Management Science, INFORMS, vol. 38(7), pages 913-925, July.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Al-Kandari Noriah M. & Lahiri Partha, 2016. "Prediction of a Function of Misclassified Binary Data," Statistics in Transition New Series, Statistics Poland, vol. 17(3), pages 429-447, September.
    2. Noriah M. Al-Kandari & Partha Lahiri, 2016. "Prediction Of A Function Of Misclassified Binary Data," Statistics in Transition New Series, Polish Statistical Association, vol. 17(3), pages 429-447, September.
    3. Rahardja, Dewi & Young, Dean M., 2010. "Credible sets for risk ratios in over-reported two-sample binomial data using the double-sampling scheme," Computational Statistics & Data Analysis, Elsevier, vol. 54(5), pages 1281-1287, May.
    4. Rahardja, Dewi & Young, Dean M., 2011. "Likelihood-based confidence intervals for the risk ratio using double sampling with over-reported binary data," Computational Statistics & Data Analysis, Elsevier, vol. 55(1), pages 813-823, January.
    5. Briceön Wiley & Chris Elrod & Phil D. Young & Dean M. Young, 2021. "An integrated‐likelihood‐ratio confidence interval for a proportion based on underreported and infallible data," Statistica Neerlandica, Netherlands Society for Statistics and Operations Research, vol. 75(3), pages 290-298, August.
    6. Dewi Rahardja, 2019. "Bayesian Inference for the Difference of Two Proportion Parameters in Over-Reported Two-Sample Binomial Data Using the Doubly Sample," Stats, MDPI, vol. 2(1), pages 1-10, February.
    7. Paul Gustafson & Nhu D. Le, 2002. "Comparing the Effects of Continuous and Discrete Covariate Mismeasurement, with Emphasis on the Dichotomization of Mismeasured Predictors," Biometrics, The International Biometric Society, vol. 58(4), pages 878-887, December.
    8. Tang, Man-Lai & Qiu, Shi-Fang & Poon, Wai-Yin, 2012. "Confidence interval construction for disease prevalence based on partial validation series," Computational Statistics & Data Analysis, Elsevier, vol. 56(5), pages 1200-1220.
    9. Dasylva Abel, 2018. "Design-Based Estimation with Record-Linked Administrative Files and a Clerical Review Sample," Journal of Official Statistics, Sciendo, vol. 34(1), pages 41-54, March.
    10. Martijn van Hasselt & Christopher R. Bollinger & Jeremy W. Bray, 2022. "A Bayesian approach to account for misclassification in prevalence and trend estimation," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 37(2), pages 351-367, March.
    11. Anil Gaba & W. Kip Viscusi, 1998. "Differences in Subjective Risk Thresholds: Worker Groups as an Example," Management Science, INFORMS, vol. 44(6), pages 801-811, June.
    12. Ben Powell & Paul A. Smith, 2020. "Computing expectations and marginal likelihoods for permutations," Computational Statistics, Springer, vol. 35(2), pages 871-891, June.
    13. Afshin Fallah & Mohsen Mohammadzadeh, 2010. "Bayesian regression analysis with linked data using mixture normal distributions," Statistical Papers, Springer, vol. 51(2), pages 421-430, June.
    14. Han Ying, 2020. "Discussion of “Small area estimation: its evolution in five decades”, by Malay Ghosh," Statistics in Transition New Series, Statistics Poland, vol. 21(4), pages 30-34, August.
    15. Durrant, Gabriele B. & D'Arrigo, Julia & Steele, Fiona, 2011. "Using field process data to predict best times of contact conditioning on household and interviewer influences," LSE Research Online Documents on Economics 52201, London School of Economics and Political Science, LSE Library.
    16. Wang Dongxu & Shen Tian & Gustafson Paul, 2012. "Partial Identification arising from Nondifferential Exposure Misclassification: How Informative are Data on the Unlikely, Maybe, and Likely Exposed?," The International Journal of Biostatistics, De Gruyter, vol. 8(1), pages 1-27, November.
    17. Kim, Gunky & Chambers, Raymond, 2012. "Regression analysis under incomplete linkage," Computational Statistics & Data Analysis, Elsevier, vol. 56(9), pages 2756-2770.
    18. Tatiana Komarova & Denis Nekipelov & Evgeny Yakovlev, 2018. "Identification, data combination, and the risk of disclosure," Quantitative Economics, Econometric Society, vol. 9(1), pages 395-440, March.
    19. Vo, Thanh Huan & Chauvet, Guillaume & Happe, André & Oger, Emmanuel & Paquelet, Stéphane & Garès, Valérie, 2023. "Extending the Fellegi-Sunter record linkage model for mixed-type data with application to the French national health data system," Computational Statistics & Data Analysis, Elsevier, vol. 179(C).
    20. Ying Han, 2020. "Discussion of "Small area estimation: its evolution in five decades", by Malay Ghosh," Statistics in Transition New Series, Polish Statistical Association, vol. 21(4), pages 30-34, August.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:csb:stintr:v:17:y:2016:i:3:p:429-447. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Beata Witek (email available below). General contact details of provider: https://edirc.repec.org/data/gusgvpl.html .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.