A machine learning approach using partitioning around medoids clustering and random forest classification to model groups of farms in regard to production parameters and bulk tank milk antibody status of two major internal parasites in dairy cows

A machine learning approach using partitioning around medoids clustering and random forest classification to model groups of farms in regard to production parameters and bulk tank milk antibody status of two major internal parasites in dairy cows

Author

Listed:

Andreas W Oehm
Andrea Springer
Daniela Jordan
Christina Strube
Gabriela Knubben-Schweizer
Katharina Charlotte Jensen
Yury Zablotski

Abstract

Fasciola hepatica and Ostertagia ostertagi are internal parasites of cattle compromising physiology, productivity, and well-being. Parasites are complex in their effect on hosts, sometimes making it difficult to identify clear directions of associations between infection and production parameters. Therefore, unsupervised approaches not assuming a structure reduce the risk of introducing bias to the analysis. They may provide insights which cannot be obtained with conventional, supervised methodology. An unsupervised, exploratory cluster analysis approach using the k–mode algorithm and partitioning around medoids detected two distinct clusters in a cross-sectional data set of milk yield, milk fat content, milk protein content as well as F. hepatica or O. ostertagi bulk tank milk antibody status from 606 dairy farms in three structurally different dairying regions in Germany. Parasite–positive farms grouped together with their respective production parameters to form separate clusters. A random forests algorithm characterised clusters with regard to external variables. Across all study regions, co–infections with F. hepatica or O. ostertagi, respectively, farming type, and pasture access appeared to be the most important factors discriminating clusters (i.e. farms). Furthermore, farm level lameness prevalence, herd size, BCS, stage of lactation, and somatic cell count were relevant criteria distinguishing clusters. This study is among the first to apply a cluster analysis approach in this context and potentially the first to implement a k–medoids algorithm and partitioning around medoids in the veterinary field. The results demonstrated that biologically relevant patterns of parasite status and milk parameters exist between farms positive for F. hepatica or O. ostertagi, respectively, and negative farms. Moreover, the machine learning approach confirmed results of previous work and shed further light on the complex setting of associations a between parasitic diseases, milk yield and milk constituents, and management practices.

Suggested Citation

Andreas W Oehm & Andrea Springer & Daniela Jordan & Christina Strube & Gabriela Knubben-Schweizer & Katharina Charlotte Jensen & Yury Zablotski, 2022. "A machine learning approach using partitioning around medoids clustering and random forest classification to model groups of farms in regard to production parameters and bulk tank milk antibody status of two major internal parasites in dairy cows," PLOS ONE, Public Library of Science, vol. 17(7), pages 1-25, July.

Handle: RePEc:plo:pone00:0271413
DOI: 10.1371/journal.pone.0271413

Download full text from publisher

References listed on IDEAS

Jerome H. Friedman & Jacqueline J. Meulman, 2004. "Clustering objects on subsets of attributes (with discussion)," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 66(4), pages 815-849, November.

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Zhaoyu Xing & Yang Wan & Juan Wen & Wei Zhong, 2024. "GOLFS: feature selection via combining both global and local information for high dimensional clustering," Computational Statistics, Springer, vol. 39(5), pages 2651-2675, July.
Jian Guo & Elizaveta Levina & George Michailidis & Ji Zhu, 2010. "Pairwise Variable Selection for High-Dimensional Model-Based Clustering," Biometrics, The International Biometric Society, vol. 66(3), pages 793-804, September.
Cathy Maugis & Gilles Celeux & Marie-Laure Martin-Magniette, 2009. "Variable Selection for Clustering with Gaussian Mixture Models," Biometrics, The International Biometric Society, vol. 65(3), pages 701-709, September.
Nicoleta Serban, 2008. "Estimating and clustering curves in the presence of heteroscedastic errors," Journal of Nonparametric Statistics, Taylor & Francis Journals, vol. 20(7), pages 553-571.
Fabio Centofanti & Antonio Lepore & Biagio Palumbo, 2024. "Sparse and smooth functional data clustering," Statistical Papers, Springer, vol. 65(2), pages 795-825, April.
Floriello, Davide & Vitelli, Valeria, 2017. "Sparse clustering of functional data," Journal of Multivariate Analysis, Elsevier, vol. 154(C), pages 1-18.
Maarten M. Kampert & Jacqueline J. Meulman & Jerome H. Friedman, 2017. "rCOSA: A Software Package for Clustering Objects on Subsets of Attributes," Journal of Classification, Springer;The Classification Society, vol. 34(3), pages 514-547, October.
Selinski, Silvia, 2006. "Similarity Measures for Clustering SNP and Epidemiological Data," Technical Reports 2006,25, Technische Universität Dortmund, Sonderforschungsbereich 475: Komplexitätsreduktion in multivariaten Datenstrukturen.
Peter D. Hoff, 2005. "Subset Clustering of Binary Sequences, with an Application to Genomic Abnormality Data," Biometrics, The International Biometric Society, vol. 61(4), pages 1027-1036, December.
Lian, Heng, 2010. "Sparse Bayesian hierarchical modeling of high-dimensional clustering problems," Journal of Multivariate Analysis, Elsevier, vol. 101(7), pages 1728-1737, August.
Gaynor, Sheila & Bair, Eric, 2017. "Identification of relevant subtypes via preweighted sparse clustering," Computational Statistics & Data Analysis, Elsevier, vol. 116(C), pages 139-154.
Nikulin, V., 2006. "Threshold-based clustering with merging and regularization in application to network intrusion detection," Computational Statistics & Data Analysis, Elsevier, vol. 51(2), pages 1184-1196, November.
Beibei Yuan & Willem Heiser & Mark Rooij, 2019. "The δ-Machine: Classification Based on Distances Towards Prototypes," Journal of Classification, Springer;The Classification Society, vol. 36(3), pages 442-470, October.
Ronglai Shen & Qianxing Mo & Nikolaus Schultz & Venkatraman E Seshan & Adam B Olshen & Jason Huse & Marc Ladanyi & Chris Sander, 2012. "Integrative Subtype Discovery in Glioblastoma Using iCluster," PLOS ONE, Public Library of Science, vol. 7(4), pages 1-9, April.
Grn, Bettina & Leisch, Friedrich, 2009. "Dealing with label switching in mixture models under genuine multimodality," Journal of Multivariate Analysis, Elsevier, vol. 100(5), pages 851-861, May.
Benhuai Xie & Wei Pan & Xiaotong Shen, 2008. "Variable Selection in Penalized Model‐Based Clustering Via Regularization on Grouped Parameters," Biometrics, The International Biometric Society, vol. 64(3), pages 921-930, September.
Francisco de A. T. Carvalho & Antonio Irpino & Rosanna Verde & Antonio Balzanella, 2022. "Batch Self-Organizing Maps for Distributional Data with an Automatic Weighting of Variables and Components," Journal of Classification, Springer;The Classification Society, vol. 39(2), pages 343-375, July.
Arias-Castro, Ery & Pu, Xiao, 2017. "A simple approach to sparse clustering," Computational Statistics & Data Analysis, Elsevier, vol. 105(C), pages 217-228.
Yang, Aijun & Jiang, Xuejun & Liu, Pengfei & Lin, Jinguan, 2016. "Sparse Bayesian multinomial probit regression model with correlation prior for high-dimensional data classification," Statistics & Probability Letters, Elsevier, vol. 119(C), pages 241-247.
Ioana Manafi & Daniela Marinescu & Monica Roman & Karen Hemming, 2017. "Mobility in Europe: Recent Trends from a Cluster Analysis," The AMFITEATRU ECONOMIC journal, Academy of Economic Studies - Bucharest, Romania, vol. 19(46), pages 711-711, August.

More about this item

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0271413. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

A machine learning approach using partitioning around medoids clustering and random forest classification to model groups of farms in regard to production parameters and bulk tank milk antibody status of two major internal parasites in dairy cows

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data