IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0353506.html

Evolving optimal text clusters: A novel GA-driven framework for dynamic ensemble fusion of multi-model contextual embeddings

Author

Listed:
  • Ali Sabah
  • Zaid Alaa

Abstract

Text clustering is an essential activity in unsupervised natural language processing (NLP), and allows the automatic structure of large-scale textual collections (in natural language processing) in news classification, routing of technical questions, and text summarisation. With the emergence of contextual embedding models, such as SBERT, RoBERTa, and DistilBERT, there has been a significant improvement in the quality of clustering, with these models producing dynamic and context-sensitive representations that are more likely to reflect domain-specific semantics than the traditional word embeddings of Word2Vec and GloVe. Although the models of contextual embedding have their own advantages, they show varying performance in different domains and the current ensemble techniques propose fixed, flat fusion weights which do not utilize their complementary abilities. There is no current solution to dynamic, label-free optimisation of weight to multi-model contextual embedding fusion in unsupervised clustering which would bridge a major gap between fixed integrative approaches and adaptive, domain-accommodative model combination. This paper presents a new Genetic Algorithm (GA)-based ensemble model that dynamically optimizes the best fusion weights of SBERT, RoBERTa, and DistilBERT embeddings without ground-truth labels. The hypothesis is that evolutionary optimisation can discover domain-adaptive weight sets that outperform standalone models as well as fixed ensemble baselines on linguistically heterogeneous datasets. The framework uses L2-normalisation and topology in which the topology is a sum-to-one constraint so that the models can be fairly integrated, and a composite fitness measure based on Silhouette Score, Adjusted Rand Index, and Topic Coherence to direct the weight evolution through tournament selection, uniform crossover, and Gaussian mutation. The proposed framework on three heterogeneous benchmarks, AG News, 20 Newsgroups, and Stack Overflow, has Silhouette Score improvements of +14–16 percent, and Topic Coherence gains of +17 percent, all statistically significant at p

Suggested Citation

  • Ali Sabah & Zaid Alaa, 2026. "Evolving optimal text clusters: A novel GA-driven framework for dynamic ensemble fusion of multi-model contextual embeddings," PLOS ONE, Public Library of Science, vol. 21(7), pages 1-25, July.
  • Handle: RePEc:plo:pone00:0353506
    DOI: 10.1371/journal.pone.0353506
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0353506
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0353506&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0353506?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0353506. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.