IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0232816.html
   My bibliography  Save this article

Solving text clustering problem using a memetic differential evolution algorithm

Author

Listed:
  • Hossam M J Mustafa
  • Masri Ayob
  • Dheeb Albashish
  • Sawsan Abu-Taleb

Abstract

The text clustering is considered as one of the most effective text document analysis methods, which is applied to cluster documents as a consequence of the expanded big data and online information. Based on the review of the related work of the text clustering algorithms, these algorithms achieved reasonable clustering results for some datasets, while they failed on a wide variety of benchmark datasets. Furthermore, the performance of these algorithms was not robust due to the inefficient balance between the exploitation and exploration capabilities of the clustering algorithm. Accordingly, this research proposes a Memetic Differential Evolution algorithm (MDETC) to solve the text clustering problem, which aims to address the effect of the hybridization between the differential evolution (DE) mutation strategy with the memetic algorithm (MA). This hybridization intends to enhance the quality of text clustering and improve the exploitation and exploration capabilities of the algorithm. Our experimental results based on six standard text clustering benchmark datasets (i.e. the Laboratory of Computational Intelligence (LABIC)) have shown that the MDETC algorithm outperformed other compared clustering algorithms based on AUC metric, F-measure, and the statistical analysis. Furthermore, the MDETC is compared with the state of art text clustering algorithms and obtained almost the best results for the standard benchmark datasets.

Suggested Citation

  • Hossam M J Mustafa & Masri Ayob & Dheeb Albashish & Sawsan Abu-Taleb, 2020. "Solving text clustering problem using a memetic differential evolution algorithm," PLOS ONE, Public Library of Science, vol. 15(6), pages 1-18, June.
  • Handle: RePEc:plo:pone00:0232816
    DOI: 10.1371/journal.pone.0232816
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0232816
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0232816&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0232816?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Mayra Z Rodriguez & Cesar H Comin & Dalcimar Casanova & Odemir M Bruno & Diego R Amancio & Luciano da F Costa & Francisco A Rodrigues, 2019. "Clustering algorithms: A comparative approach," PLOS ONE, Public Library of Science, vol. 14(1), pages 1-34, January.
    2. Hossam M J Mustafa & Masri Ayob & Mohd Zakree Ahmad Nazri & Graham Kendall, 2019. "An improved adaptive memetic differential evolution optimization algorithms for data clustering problems," PLOS ONE, Public Library of Science, vol. 14(5), pages 1-28, May.
    3. Rahab M Ramadan & Safa M Gasser & Mohamed S El-Mahallawy & Karim Hammad & Ahmed M El Bakly, 2018. "A memetic optimization algorithm for multi-constrained multicast routing in ad hoc networks," PLOS ONE, Public Library of Science, vol. 13(3), pages 1-17, March.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Hossam M J Mustafa & Masri Ayob & Mohd Zakree Ahmad Nazri & Graham Kendall, 2019. "An improved adaptive memetic differential evolution optimization algorithms for data clustering problems," PLOS ONE, Public Library of Science, vol. 14(5), pages 1-28, May.
    2. Fernandez Martinez, Roberto & Lostado Lorza, Ruben & Santos Delgado, Ana Alexandra & Piedra, Nelson, 2021. "Use of classification trees and rule-based models to optimize the funding assignment to research projects: A case study of UTPL," Journal of Informetrics, Elsevier, vol. 15(1).
    3. Corrêa, Edilson A. & Marinho, Vanessa Q. & Amancio, Diego R., 2020. "Semantic flow in language networks discriminates texts by genre and publication date," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 557(C).
    4. Ebba Mark & Ryan Rafaty & Moritz Schwarz, 2022. "Spatial-temporal dynamics of employment shocks in declining coal mining regions and potentialities of the 'just transition'," Papers 2211.12619, arXiv.org.
    5. Simon Crase & Suresh N Thennadil, 2022. "An analysis framework for clustering algorithm selection with applications to spectroscopy," PLOS ONE, Public Library of Science, vol. 17(3), pages 1-24, March.
    6. K. S. Sablin & E. S. Kagan & E. S. Chernova, 2020. "Clustering of the Russian coal mining regions: Investment and innovation activity," Journal of New Economy, Ural State University of Economics, vol. 21(1), pages 89-106, March.
    7. Narjes Vara & Mahdieh Mirzabeigi & Hajar Sotudeh & Seyed Mostafa Fakhrahmad, 2022. "Application of k-means clustering algorithm to improve effectiveness of the results recommended by journal recommender system," Scientometrics, Springer;Akadémiai Kiadó, vol. 127(6), pages 3237-3252, June.
    8. Alfred Kume & Stephen G Walker, 2021. "The utility of clusters and a Hungarian clustering algorithm," PLOS ONE, Public Library of Science, vol. 16(8), pages 1-23, August.
    9. Quispe, Laura V.C. & Tohalino, Jorge A.V. & Amancio, Diego R., 2021. "Using virtual edges to improve the discriminability of co-occurrence text networks," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 562(C).
    10. Sultan Mahmud & Ferdausi Mahojabin Sumana & Md Mohsin & Md. Hasinur Rahaman Khan, 2022. "Redefining homogeneous climate regions in Bangladesh using multivariate clustering approaches," Natural Hazards: Journal of the International Society for the Prevention and Mitigation of Natural Hazards, Springer;International Society for the Prevention and Mitigation of Natural Hazards, vol. 111(2), pages 1863-1884, March.
    11. Tohalino, Jorge A.V. & Amancio, Diego R., 2022. "On predicting research grants productivity via machine learning," Journal of Informetrics, Elsevier, vol. 16(2).
    12. Christian Hennig, 2022. "An empirical comparison and characterisation of nine popular clustering methods," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 16(1), pages 201-229, March.
    13. Mikhail Kanevski, 2021. "Unsupervised learning of Swiss population spatial distribution," PLOS ONE, Public Library of Science, vol. 16(2), pages 1-24, February.
    14. Trotta, Gianluca, 2020. "An empirical analysis of domestic electricity load profiles: Who consumes how much and when?," Applied Energy, Elsevier, vol. 275(C).

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0232816. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.