IDEAS home Printed from https://ideas.repec.org/a/bla/stanee/v76y2022i1p65-96.html
   My bibliography  Save this article

Robust prediction of domain compositions from uncertain data using isometric logratio transformations in a penalized multivariate Fay–Herriot model

Author

Listed:
  • Joscha Krause
  • Jan Pablo Burgard
  • Domingo Morales

Abstract

Assessing regional population compositions is an important task in many research fields. Small area estimation with generalized linear mixed models marks a powerful tool for this purpose. However, the method has limitations in practice. When the data are subject to measurement errors, small area models produce inefficient or biased results since they cannot account for data uncertainty. This is particularly problematic for composition prediction, since generalized linear mixed models often rely on approximate likelihood inference. Obtained predictions are not reliable. We propose a robust multivariate Fay–Herriot model to solve these issues. It combines compositional data analysis with robust optimization theory. The nonlinear estimation of compositions is restated as a linear problem through isometric logratio transformations. Robust model parameter estimation is performed via penalized maximum likelihood. A robust best predictor is derived. Simulations are conducted to demonstrate the effectiveness of the approach. An application to alcohol consumption in Germany is provided.

Suggested Citation

  • Joscha Krause & Jan Pablo Burgard & Domingo Morales, 2022. "Robust prediction of domain compositions from uncertain data using isometric logratio transformations in a penalized multivariate Fay–Herriot model," Statistica Neerlandica, Netherlands Society for Statistics and Operations Research, vol. 76(1), pages 65-96, February.
  • Handle: RePEc:bla:stanee:v:76:y:2022:i:1:p:65-96
    DOI: 10.1111/stan.12253
    as

    Download full text from publisher

    File URL: https://doi.org/10.1111/stan.12253
    Download Restriction: no

    File URL: https://libkey.io/10.1111/stan.12253?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Wang, Huiwen & Liu, Qiang & Mok, Henry M.K. & Fu, Linghui & Tse, Wai Man, 2007. "A hyperspherical transformation forecasting model for compositional data," European Journal of Operational Research, Elsevier, vol. 179(2), pages 459-468, June.
    2. Shonosuke Sugasawa & Hiromasa Tamae & Tatsuya Kubokawa, 2017. "Bayesian Estimators for Small Area Models Shrinking Both Means and Variances," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 44(1), pages 150-167, March.
    3. Bertsimas, Dimitris & Copenhaver, Martin S., 2018. "Characterization of the equivalence of robustification and regularization in linear and matrix regression," European Journal of Operational Research, Elsevier, vol. 270(3), pages 931-942.
    4. Tapabrata Maiti & Hao Ren & Samiran Sinha, 2014. "Prediction Error of Small Area Predictors Shrinking Both Means and Variances," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 41(3), pages 775-790, September.
    5. Esther López-Vizcaíno & María José Lombardía & Domingo Morales, 2015. "Small area estimation of labour force indicators under a multinomial model with correlated time and area effects," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 178(3), pages 535-565, June.
    6. Joanna Morais & Christine Thomas-Agnan & Michel Simioni, 2017. "Interpretation of explanatory variables impacts in compositional regression models," Working Papers hal-01563362, HAL.
    7. Juan José Egozcue & Vera Pawlowsky-Glahn, 2019. "Compositional data: the sample space and its structure," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(3), pages 599-638, September.
    8. Lynn M. R. Ybarra & Sharon L. Lohr, 2008. "Small area estimation when auxiliary information is measured with error," Biometrika, Biometrika Trust, vol. 95(4), pages 919-931.
    9. Tomáš Hobza & Domingo Morales & Laureano Santamaría, 2018. "Small area estimation of poverty proportions under unit-level temporal binomial-logit mixed models," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 27(2), pages 270-294, June.
    10. J. L. Scealy & A. H. Welsh, 2017. "A Directional Mixed Effects Model for Compositional Expenditure Data," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 112(517), pages 24-36, January.
    11. Friedman, Jerome H. & Hastie, Trevor & Tibshirani, Rob, 2010. "Regularization Paths for Generalized Linear Models via Coordinate Descent," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 33(i01).
    12. Roberto Benavent & Domingo Morales, 2021. "Small area estimation under a temporal bivariate area-level linear mixed model with independent time effects," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 30(1), pages 195-222, March.
    13. Azka Ubaidillah & Khairil Anwar Notodiputro & Anang Kurnia & I. Wayan Mangku, 2019. "Multivariate Fay-Herriot models for small area estimation with application to household consumption per capita expenditure in Indonesia," Journal of Applied Statistics, Taylor & Francis Journals, vol. 46(15), pages 2845-2861, November.
    14. Isabel Molina & Ayoub Saei & M. José Lombardía, 2007. "Small area estimates of labour force participation under a multinomial logit mixed model," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 170(4), pages 975-1000, October.
    15. A. F. Militino & M. D. Ugarte & T. Goicoa, 2015. "Deriving small area estimates from information technology business surveys," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 178(4), pages 1051-1067, October.
    16. Crum, R.M. & Helzer, J.E. & Anthony, J.C., 1993. "Level of education and alcohol abuse and dependence in adulthood: A further inquiry," American Journal of Public Health, American Public Health Association, vol. 83(6), pages 830-837.
    17. Juan José Egozcue & Vera Pawlowsky-Glahn, 2019. "Rejoinder on: Compositional data: the sample space and its structure," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(3), pages 658-663, September.
    18. Serena Arima & William R. Bell & Gauri S. Datta & Carolina Franco & Brunero Liseo, 2017. "Multivariate Fay–Herriot Bayesian estimation of small area means under functional measurement error," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 180(4), pages 1191-1209, October.
    19. Hui Zou & Trevor Hastie, 2005. "Addendum: Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(5), pages 768-768, November.
    20. Mills, Terence C., 2018. "Is There Convergence in National Alcohol Consumption Patterns? Evidence from a Compositional Time Series Approach," Journal of Wine Economics, Cambridge University Press, vol. 13(1), pages 92-98, February.
    21. Gert G. Wagner & Joachim R. Frick & Jürgen Schupp, 2007. "The German Socio-Economic Panel Study (SOEP) – Scope, Evolution and Enhancements," Schmollers Jahrbuch : Journal of Applied Social Science Studies / Zeitschrift für Wirtschafts- und Sozialwissenschaften, Duncker & Humblot, Berlin, vol. 127(1), pages 139-169.
    22. Ray Chambers & Nicola Salvati & Nikos Tzavidis, 2016. "Semiparametric small area estimation for binary outcomes with application to unemployment estimation for local authorities in the UK," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 179(2), pages 453-479, February.
    23. Hui Zou & Trevor Hastie, 2005. "Regularization and variable selection via the elastic net," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 67(2), pages 301-320, April.
    24. Hart, Jarrett & Alston, Julian M., 2019. "Persistent Patterns in the U.S. Alcohol Market: Looking at the Link between Demographics and Drinking," Journal of Wine Economics, Cambridge University Press, vol. 14(4), pages 356-364, November.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. María Dolores Esteban & María José Lombardía & Esther López-Vizcaíno & Domingo Morales & Agustín Pérez, 2023. "Small area estimation of average compositions under multivariate nested error regression models," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 32(2), pages 651-676, June.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. María Dolores Esteban & María José Lombardía & Esther López-Vizcaíno & Domingo Morales & Agustín Pérez, 2023. "Small area estimation of average compositions under multivariate nested error regression models," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 32(2), pages 651-676, June.
    2. Domingo Morales & Joscha Krause & Jan Pablo Burgard, 2022. "On the Use of Aggregate Survey Data for Estimating Regional Major Depressive Disorder Prevalence," Psychometrika, Springer;The Psychometric Society, vol. 87(1), pages 344-368, March.
    3. María Dolores Esteban & María José Lombardía & Esther López-Vizcaíno & Domingo Morales & Agustín Pérez, 2020. "Small area estimation of proportions under area-level compositional mixed models," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 29(3), pages 793-818, September.
    4. María Dolores Esteban & María José Lombardía & Esther López‐Vizcaíno & Domingo Morales & Agustín Pérez, 2022. "Empirical best prediction of small area bivariate parameters," Scandinavian Journal of Statistics, Danish Society for Theoretical Statistics;Finnish Statistical Society;Norwegian Statistical Association;Swedish Statistical Association, vol. 49(4), pages 1699-1727, December.
    5. Jan Pablo Burgard & Joscha Krause & Dennis Kreber, 2019. "Regularized Area-level Modelling for Robust Small Area Estimation in the Presence of Unknown Covariate Measurement Errors," Research Papers in Economics 2019-04, University of Trier, Department of Economics.
    6. Jan Pablo Burgard & Joscha Krause & Ralf Münnich, 2019. "Penalized Small Area Models for the Combination of Unit- and Area-level Data," Research Papers in Economics 2019-05, University of Trier, Department of Economics.
    7. Joscha Krause & Jan Pablo Burgard & Domingo Morales, 2022. "$$\ell _2$$ ℓ 2 -penalized approximate likelihood inference in logit mixed models for regional prevalence estimation under covariate rank-deficiency," Metrika: International Journal for Theoretical and Applied Statistics, Springer, vol. 85(4), pages 459-489, May.
    8. Jan Pablo Burgard & Domingo Morales & Anna-Lena Wölwer, 2022. "Small area estimation of socioeconomic indicators for sampled and unsampled domains," AStA Advances in Statistical Analysis, Springer;German Statistical Society, vol. 106(2), pages 287-314, June.
    9. Angelo Moretti, 2023. "Estimation of small area proportions under a bivariate logistic mixed model," Quality & Quantity: International Journal of Methodology, Springer, vol. 57(4), pages 3663-3684, August.
    10. Jan Pablo Burgard & Joscha Krause & Dennis Kreber & Domingo Morales, 2021. "The generalized equivalence of regularization and min–max robustification in linear mixed models," Statistical Papers, Springer, vol. 62(6), pages 2857-2883, December.
    11. Jan Pablo Burgard & Joscha Krause & Domingo Morales, 2022. "A measurement error Rao–Yu model for regional prevalence estimation over time using uncertain data obtained from dependent survey estimates," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 31(1), pages 204-234, March.
    12. Jan Pablo Burgard & María Dolores Esteban & Domingo Morales & Agustín Pérez, 2021. "Small area estimation under a measurement error bivariate Fay–Herriot model," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 30(1), pages 79-108, March.
    13. Tutz, Gerhard & Pößnecker, Wolfgang & Uhlmann, Lorenz, 2015. "Variable selection in general multinomial logit models," Computational Statistics & Data Analysis, Elsevier, vol. 82(C), pages 207-222.
    14. Mkhadri, Abdallah & Ouhourane, Mohamed, 2013. "An extended variable inclusion and shrinkage algorithm for correlated variables," Computational Statistics & Data Analysis, Elsevier, vol. 57(1), pages 631-644.
    15. Susan Athey & Guido W. Imbens & Stefan Wager, 2018. "Approximate residual balancing: debiased inference of average treatment effects in high dimensions," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 80(4), pages 597-623, September.
    16. Christopher J Greenwood & George J Youssef & Primrose Letcher & Jacqui A Macdonald & Lauryn J Hagg & Ann Sanson & Jenn Mcintosh & Delyse M Hutchinson & John W Toumbourou & Matthew Fuller-Tyszkiewicz &, 2020. "A comparison of penalised regression methods for informing the selection of predictive markers," PLOS ONE, Public Library of Science, vol. 15(11), pages 1-14, November.
    17. Immanuel Bayer & Philip Groth & Sebastian Schneckener, 2013. "Prediction Errors in Learning Drug Response from Gene Expression Data – Influence of Labeling, Sample Size, and Machine Learning Algorithm," PLOS ONE, Public Library of Science, vol. 8(7), pages 1-13, July.
    18. Mostafa Rezaei & Ivor Cribben & Michele Samorani, 2021. "A clustering-based feature selection method for automatically generated relational attributes," Annals of Operations Research, Springer, vol. 303(1), pages 233-263, August.
    19. Gustavo A. Alonso-Silverio & Víctor Francisco-García & Iris P. Guzmán-Guzmán & Elías Ventura-Molina & Antonio Alarcón-Paredes, 2021. "Toward Non-Invasive Estimation of Blood Glucose Concentration: A Comparative Performance," Mathematics, MDPI, vol. 9(20), pages 1-13, October.
    20. Christopher Kath & Florian Ziel, 2018. "The value of forecasts: Quantifying the economic gains of accurate quarter-hourly electricity price forecasts," Papers 1811.08604, arXiv.org.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bla:stanee:v:76:y:2022:i:1:p:65-96. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Wiley Content Delivery (email available below). General contact details of provider: http://www.blackwellpublishing.com/journal.asp?ref=0039-0402 .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.