IDEAS home Printed from https://ideas.repec.org/a/taf/japsta/v44y2017i4p734-752.html
   My bibliography  Save this article

Exploratory tools for outlier detection in compositional data with structural zeros

Author

Listed:
  • M. Templ
  • K. Hron
  • P. Filzmoser

Abstract

The analysis of compositional data using the log-ratio approach is based on ratios between the compositional parts. Zeros in the parts thus cause serious difficulties for the analysis. This is a particular problem in case of structural zeros, which cannot be simply replaced by a non-zero value as it is done, e.g. for values below detection limit or missing values. Instead, zeros to be incorporated into further statistical processing. The focus is on exploratory tools for identifying outliers in compositional data sets with structural zeros. For this purpose, Mahalanobis distances are estimated, computed either directly for subcompositions determined by their zero patterns, or by using imputation to improve the efficiency of the estimates, and then proceed to the subcompositional and subgroup level. For this approach, new theory is formulated that allows to estimate covariances for imputed compositional data and to apply estimations on subgroups using parts of this covariance matrix. Moreover, the zero pattern structure is analyzed using principal component analysis for binary data to achieve a comprehensive view of the overall multivariate data structure. The proposed tools are applied to larger compositional data sets from official statistics, where the need for an appropriate treatment of zeros is obvious.

Suggested Citation

  • M. Templ & K. Hron & P. Filzmoser, 2017. "Exploratory tools for outlier detection in compositional data with structural zeros," Journal of Applied Statistics, Taylor & Francis Journals, vol. 44(4), pages 734-752, March.
  • Handle: RePEc:taf:japsta:v:44:y:2017:i:4:p:734-752
    DOI: 10.1080/02664763.2016.1182135
    as

    Download full text from publisher

    File URL: http://hdl.handle.net/10.1080/02664763.2016.1182135
    Download Restriction: Access to full text is restricted to subscribers.

    File URL: https://libkey.io/10.1080/02664763.2016.1182135?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Martín-Fernández, J.A. & Hron, K. & Templ, M. & Filzmoser, P. & Palarea-Albaladejo, J., 2012. "Model-based replacement of rounded zeros in compositional data: Classical and robust approaches," Computational Statistics & Data Analysis, Elsevier, vol. 56(9), pages 2688-2704.
    2. de Leeuw, Jan, 2006. "Principal component analysis of binary data by iterated singular value decomposition," Computational Statistics & Data Analysis, Elsevier, vol. 50(1), pages 21-39, January.
    3. J. L. Scealy & A. H. Welsh, 2011. "Regression for compositional data by using distributions defined on the hypersphere," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 73(3), pages 351-375, June.
    4. Hron, K. & Templ, M. & Filzmoser, P., 2010. "Imputation of missing values for compositional data using classical and robust methods," Computational Statistics & Data Analysis, Elsevier, vol. 54(12), pages 3095-3107, December.
    5. Valentin Todorov & Matthias Templ & Peter Filzmoser, 2011. "Detection of multivariate outliers in business survey data with incomplete information," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 5(1), pages 37-56, April.
    6. Andreas Alfons & Stefan Kraft & Matthias Templ & Peter Filzmoser, 2011. "Simulation of close-to-reality population data for household surveys with application to EU-SILC," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 20(3), pages 383-407, August.
    7. Wang, Huiwen & Liu, Qiang & Mok, Henry M.K. & Fu, Linghui & Tse, Wai Man, 2007. "A hyperspherical transformation forecasting model for compositional data," European Journal of Operational Research, Elsevier, vol. 179(2), pages 459-468, June.
    8. Jane Fry & Tim Fry & Keith McLaren, 2000. "Compositional data analysis and zeros in micro data," Applied Economics, Taylor & Francis Journals, vol. 32(8), pages 953-959.
    9. Matthias Templ & Andreas Alfons & Peter Filzmoser, 2012. "Exploring incomplete data using visualization techniques," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 6(1), pages 29-47, April.
    10. Alfons, Andreas & Templ, Matthias, 2013. "Estimation of Social Exclusion Indicators from Complex Surveys: The R Package laeken," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 54(i15).
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Farnè, Matteo & Vouldis, Angelos T., 2018. "A methodology for automised outlier detection in high-dimensional datasets: an application to euro area banks' supervisory data," Working Paper Series 2171, European Central Bank.
    2. Dorothea Dumuid & Željko Pedišić & Javier Palarea-Albaladejo & Josep Antoni Martín-Fernández & Karel Hron & Timothy Olds, 2020. "Compositional Data Analysis in Time-Use Epidemiology: What, Why, How," IJERPH, MDPI, vol. 17(7), pages 1-17, March.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Tsagris, Michail, 2014. "The k-NN algorithm for compositional data: a revised approach with and without zero values present," MPRA Paper 65866, University Library of Munich, Germany.
    2. Juan José Egozcue & Vera Pawlowsky-Glahn, 2019. "Compositional data: the sample space and its structure," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(3), pages 599-638, September.
    3. Templ, Matthias & Meindl, Bernhard & Kowarik, Alexander & Dupriez, Olivier, 2017. "Simulation of Synthetic Complex Data: The R Package simPop," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 79(i10).
    4. Templ Matthias, 2015. "Quality Indicators for Statistical Disclosure Methods: A Case Study on the Structure of Earnings Survey," Journal of Official Statistics, Sciendo, vol. 31(4), pages 737-761, December.
    5. Tsagris, Michail & Preston, Simon & T.A. Wood, Andrew, 2016. "Improved classi cation for compositional data using the $\alpha$-transformation," MPRA Paper 67657, University Library of Munich, Germany.
    6. Morais, Joanna & Simioni, Michel & Thomas-Agnan, Christine, 2016. "A tour of regression models for explaining shares," TSE Working Papers 16-742, Toulouse School of Economics (TSE).
    7. Michail Tsagris & Simon Preston & Andrew T. A. Wood, 2016. "Improved Classification for Compositional Data Using the α-transformation," Journal of Classification, Springer;The Classification Society, vol. 33(2), pages 243-261, July.
    8. Matthias Templ & Andreas Alfons & Peter Filzmoser, 2012. "Exploring incomplete data using visualization techniques," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 6(1), pages 29-47, April.
    9. Alfons, Andreas & Templ, Matthias, 2013. "Estimation of Social Exclusion Indicators from Complex Surveys: The R Package laeken," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 54(i15).
    10. Takahiro Yoshida & Morito Tsutsumi, 2018. "On the effects of spatial relationships in spatial compositional multivariate models," Letters in Spatial and Resource Sciences, Springer, vol. 11(1), pages 57-70, March.
    11. Maria Anna Di Palma & Michele Gallo, 2019. "External Information Model in a Compositional Perspective: Evaluation of Campania Adolescents’ Preferences in the Allocation of Leisure-Time," Social Indicators Research: An International and Interdisciplinary Journal for Quality-of-Life Measurement, Springer, vol. 146(1), pages 117-133, November.
    12. Templ, Matthias & Kowarik, Alexander & Meindl, Bernhard, 2015. "Statistical Disclosure Control for Micro-Data Using the R Package sdcMicro," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 67(i04).
    13. Elena Catanese, 2016. "Data Editing for Complex Surveys in Presence Of Administrative Data: An Application to Fss 2013 Livestock Survey Data Based on The Joint Sequential Use Of Different R Packages," Romanian Statistical Review, Romanian Statistical Review, vol. 64(2), pages 101-117, June.
    14. Kowarik, Alexander & Templ, Matthias, 2016. "Imputation with the R Package VIM," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 74(i07).
    15. Terence C. Mills, 2009. "Forecasting obesity trends in England," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 172(1), pages 107-117, January.
    16. Matthias Templ & Alexander Kowarik & Bernhard Meindl, 2014. "Development and Current Practice in Using R at Statistics Austria," Romanian Statistical Review, Romanian Statistical Review, vol. 62(2), pages 173-184, June.
    17. Tsagris, Michail, 2015. "Regression analysis with compositional data containing zero values," MPRA Paper 67868, University Library of Munich, Germany.
    18. Nikola Štefelová & Andreas Alfons & Javier Palarea-Albaladejo & Peter Filzmoser & Karel Hron, 2021. "Robust regression with compositional covariates including cellwise outliers," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 15(4), pages 869-909, December.
    19. Wang, Fa, 2017. "Maximum likelihood estimation and inference for high dimensional nonlinear factor models with application to factor-augmented regressions," MPRA Paper 93484, University Library of Munich, Germany, revised 19 May 2019.
    20. Patrick Krennmair & Timo Schmid, 2022. "Flexible domain prediction using mixed effects random forests," Journal of the Royal Statistical Society Series C, Royal Statistical Society, vol. 71(5), pages 1865-1894, November.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:taf:japsta:v:44:y:2017:i:4:p:734-752. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Chris Longhurst (email available below). General contact details of provider: http://www.tandfonline.com/CJAS20 .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.