IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v128y2018icp184-199.html
   My bibliography  Save this article

ICS for multivariate outlier detection with application to quality control

Author

Listed:
  • Archimbaud, Aurore
  • Nordhausen, Klaus
  • Ruiz-Gazen, Anne

Abstract

In high reliability standards fields such as automotive, avionics or aerospace, the detection of anomalies is crucial. An efficient methodology for automatically detecting multivariate outliers is introduced. It takes advantage of the remarkable properties of the Invariant Coordinate Selection (ICS) method which leads to an affine invariant coordinate system in which the Euclidian distance corresponds to a Mahalanobis Distance (MD) in the original coordinates. The limitations of MD are highlighted using theoretical arguments in a context where the dimension of the data is large. Owing to the resulting dimension reduction, ICS is expected to improve the power of outlier detection rules such as MD-based criteria. The paper includes practical guidelines for using ICS in the context of a small proportion of outliers. The use of the regular covariance matrix and the so called matrix of fourth moments as the scatter pair is recommended. This choice combines the simplicity of implementation together with the possibility to derive theoretical results. The selection of relevant invariant components through parallel analysis and normality tests is addressed. A simulation study confirms the good properties of the proposal and provides a comparison with Principal Component Analysis and MD. The performance of the proposal is also evaluated on two real data sets using a user-friendly R package accompanying the paper.

Suggested Citation

  • Archimbaud, Aurore & Nordhausen, Klaus & Ruiz-Gazen, Anne, 2018. "ICS for multivariate outlier detection with application to quality control," Computational Statistics & Data Analysis, Elsevier, vol. 128(C), pages 184-199.
  • Handle: RePEc:eee:csdana:v:128:y:2018:i:c:p:184-199
    DOI: 10.1016/j.csda.2018.06.011
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167947318301579
    Download Restriction: Full text for ScienceDirect subscribers only.

    File URL: https://libkey.io/10.1016/j.csda.2018.06.011?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Alashwali, Fatimah & Kent, John T., 2016. "The use of a common location measure in the invariant coordinate selection and projection pursuit," Journal of Multivariate Analysis, Elsevier, vol. 152(C), pages 145-161.
    2. Claudio Agostinelli & Andy Leung & Victor Yohai & Ruben Zamar, 2015. "Rejoinder on: Robust estimation of multivariate location and scatter in the presence of cellwise and casewise contamination," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 24(3), pages 484-488, September.
    3. Cerioli, Andrea & Farcomeni, Alessio, 2011. "Error rates for multivariate outlier detection," Computational Statistics & Data Analysis, Elsevier, vol. 55(1), pages 544-553, January.
    4. Li, Baibing & Martin, Elaine B. & Morris, A. Julian, 2002. "On principal component analysis in L1," Computational Statistics & Data Analysis, Elsevier, vol. 40(3), pages 471-474, September.
    5. Nordhausen, Klaus & Oja, Hannu & Tyler, David E., 2008. "Tools for Exploring Multivariate Data: The Package ICS," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 28(i06).
    6. Dray, Stephane, 2008. "On the number of principal components: A test of dimensionality based on measurements of similarity between matrices," Computational Statistics & Data Analysis, Elsevier, vol. 52(4), pages 2228-2237, January.
    7. David E. Tyler & Frank Critchley & Lutz Dümbgen & Hannu Oja, 2009. "Invariant co‐ordinate selection," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 71(3), pages 549-592, June.
    8. Caussinus, H. & Fekri, M. & Hakam, S. & Ruiz-Gazen, A., 2003. "A monitoring display of multivariate outliers," Computational Statistics & Data Analysis, Elsevier, vol. 44(1-2), pages 237-252, October.
    9. Bonett, Douglas G. & Seier, Edith, 2002. "A test of normality with high uniform power," Computational Statistics & Data Analysis, Elsevier, vol. 40(3), pages 435-445, September.
    10. Claudio Agostinelli & Andy Leung & Victor Yohai & Ruben Zamar, 2015. "Robust estimation of multivariate location and scatter in the presence of cellwise and casewise contamination," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 24(3), pages 441-461, September.
    11. Peres-Neto, Pedro R. & Jackson, Donald A. & Somers, Keith M., 2005. "How many principal components? stopping rules for determining the number of non-trivial axes revisited," Computational Statistics & Data Analysis, Elsevier, vol. 49(4), pages 974-997, June.
    12. Klaus Nordhausen & David E. Tyler, 2015. "A cautionary note on robust covariance plug-in methods," Biometrika, Biometrika Trust, vol. 102(3), pages 573-588.
    13. Croux, Christophe & Haesbroeck, Gentiane, 1999. "Influence Function and Efficiency of the Minimum Covariance Determinant Scatter Matrix Estimator," Journal of Multivariate Analysis, Elsevier, vol. 71(2), pages 161-190, November.
    14. Cerioli, Andrea, 2010. "Multivariate Outlier Detection With High-Breakdown Estimators," Journal of the American Statistical Association, American Statistical Association, vol. 105(489), pages 147-156.
    15. Todorov, Valentin & Filzmoser, Peter, 2009. "An Object-Oriented Framework for Robust Multivariate Analysis," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 32(i03).
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Loperfido, Nicola, 2021. "Some theoretical properties of two kurtosis matrices, with application to invariant coordinate selection," Journal of Multivariate Analysis, Elsevier, vol. 186(C).
    2. Ruiz-Gazen, Anne & Thomas-Agnan, Christine & Laurent, Thibault & Mondon, Camille, 2022. "Detecting outliers in compositional data using Invariant Coordinate Selection," TSE Working Papers 22-1320, Toulouse School of Economics (TSE).
    3. Klaus Nordhausen & Anne Ruiz-Gazen, 2022. "On the usage of joint diagonalization in multivariate statistics," Post-Print hal-04296111, HAL.
    4. Archimbaud, Aurore & Boulfani, Fériel & Gendre, Xavier & Nordhausen, Klaus & Ruiz-Gazen, Anne & Virta, Joni, 2021. "ICS for multivariate functional anomaly detection with applications to predictive maintenance and quality control," TSE Working Papers 21-1182, Toulouse School of Economics (TSE), revised Mar 2022.
    5. Nordhausen, Klaus & Ruiz-Gazen, Anne, 2021. "On the usage of joint diagonalization in multivariate statistics," TSE Working Papers 21-1268, Toulouse School of Economics (TSE).
    6. Nordhausen, Klaus & Ruiz-Gazen, Anne, 2022. "On the usage of joint diagonalization in multivariate statistics," Journal of Multivariate Analysis, Elsevier, vol. 188(C).

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Nordhausen, Klaus & Ruiz-Gazen, Anne, 2022. "On the usage of joint diagonalization in multivariate statistics," Journal of Multivariate Analysis, Elsevier, vol. 188(C).
    2. Jan Kalina & Jan Tichavský, 2022. "The minimum weighted covariance determinant estimator for high-dimensional data," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 16(4), pages 977-999, December.
    3. Klaus Nordhausen & Anne Ruiz-Gazen, 2022. "On the usage of joint diagonalization in multivariate statistics," Post-Print hal-04296111, HAL.
    4. Nordhausen, Klaus & Ruiz-Gazen, Anne, 2021. "On the usage of joint diagonalization in multivariate statistics," TSE Working Papers 21-1268, Toulouse School of Economics (TSE).
    5. Cerioli, Andrea & Farcomeni, Alessio & Riani, Marco, 2013. "Robust distances for outlier-free goodness-of-fit testing," Computational Statistics & Data Analysis, Elsevier, vol. 65(C), pages 29-45.
    6. Nordhausen, Klaus & Oja, Hannu & Tyler, David E., 2022. "Asymptotic and bootstrap tests for subspace dimension," Journal of Multivariate Analysis, Elsevier, vol. 188(C).
    7. Alashwali, Fatimah & Kent, John T., 2016. "The use of a common location measure in the invariant coordinate selection and projection pursuit," Journal of Multivariate Analysis, Elsevier, vol. 152(C), pages 145-161.
    8. Marco Riani & Andrea Cerioli & Francesca Torti, 2014. "On consistency factors and efficiency of robust S-estimators," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 23(2), pages 356-387, June.
    9. Dümbgen, Lutz & Nordhausen, Klaus & Schuhmacher, Heike, 2016. "New algorithms for M-estimation of multivariate scatter and location," Journal of Multivariate Analysis, Elsevier, vol. 144(C), pages 200-217.
    10. Maronna, Ricardo A. & Yohai, Victor J., 2017. "Robust and efficient estimation of multivariate scatter and location," Computational Statistics & Data Analysis, Elsevier, vol. 109(C), pages 64-75.
    11. Henry Velasco & Henry Laniado & Mauricio Toro & Víctor Leiva & Yuhlong Lio, 2020. "Robust Three-Step Regression Based on Comedian and Its Performance in Cell-Wise and Case-Wise Outliers," Mathematics, MDPI, vol. 8(8), pages 1-18, August.
    12. Silvia Salini & Andrea Cerioli & Fabrizio Laurini & Marco Riani, 2016. "Reliable Robust Regression Diagnostics," International Statistical Review, International Statistical Institute, vol. 84(1), pages 99-127, April.
    13. Josse, Julie & Husson, François, 2012. "Selecting the number of components in principal component analysis using cross-validation approximations," Computational Statistics & Data Analysis, Elsevier, vol. 56(6), pages 1869-1879.
    14. Kelly P. Murillo & Eugenio M. Rocha, 2020. "Factors Influencing the Economic Behavior of the Food, Beverages and Tobacco Industry: A Case Study for Portuguese Enterprises," World Journal of Applied Economics, WERI-World Economic Research Institute, vol. 6(2), pages 99-121, December.
    15. Andrea Cerioli & Marco Riani & Anthony C. Atkinson & Aldo Corbellini, 2018. "The power of monitoring: how to make the most of a contaminated multivariate sample," Statistical Methods & Applications, Springer;Società Italiana di Statistica, vol. 27(4), pages 559-587, December.
    16. Leung, Andy & Zhang, Hongyang & Zamar, Ruben, 2016. "Robust regression estimation and inference in the presence of cellwise and casewise contamination," Computational Statistics & Data Analysis, Elsevier, vol. 99(C), pages 1-11.
    17. Sergio Camiz & Valério D. Pillar, 2018. "Identifying the Informational/Signal Dimension in Principal Component Analysis," Mathematics, MDPI, vol. 6(11), pages 1-16, November.
    18. Claudio Agostinelli & Luca Greco, 2019. "Weighted likelihood estimation of multivariate location and scatter," TEST: An Official Journal of the Spanish Society of Statistics and Operations Research, Springer;Sociedad de Estadística e Investigación Operativa, vol. 28(3), pages 756-784, September.
    19. Nikola Štefelová & Andreas Alfons & Javier Palarea-Albaladejo & Peter Filzmoser & Karel Hron, 2021. "Robust regression with compositional covariates including cellwise outliers," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 15(4), pages 869-909, December.
    20. Joy R. Petway & Yu-Pin Lin & Rainer F. Wunderlich, 2019. "Analyzing Opinions on Sustainable Agriculture: Toward Increasing Farmer Knowledge of Organic Practices in Taiwan-Yuanli Township," Sustainability, MDPI, vol. 11(14), pages 1-27, July.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:128:y:2018:i:c:p:184-199. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.