IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0354718.html

Contrastive learning with mutual information enhancement and negative sample augmentation combined with KAN for text clustering

Author

Listed:
  • Yuanmin Zhang
  • Hao Li
  • Chunzhi Xie
  • Yong Huang
  • Zhenyi Wu
  • Yanjun Li

Abstract

This paper proposes a text clustering model based on mutual information-enhanced contrastive learning combined with Kolmogorov-Arnold Networks (KAN) to address challenges in text clustering, including insufficient robustness of text representations, feature redundancy, and the curse of dimensionality. Text clustering is essential for organizing unstructured textual data, yet existing methods often suffer from weak feature representations and limited scalability in high-dimensional spaces. The proposed model jointly addresses these three core issues by maximizing mutual information to capture nonlinear relationships among positive samples, expanding negative samples to enhance discriminability, and employing the KAN architecture to adapt to the complex structures of high-dimensional data. Experimental results on eight benchmark datasets demonstrate that the proposed model achieves state-of-the-art accuracy on seven out of eight datasets and leads normalized mutual information on six datasets, outperforming existing text clustering baselines. To validate its practical utility, the model was applied to cluster Weibo posts collected via web crawlers using “technology” as the keyword during January and February 2025. The case study results reveal that the model effectively identifies meaningful thematic clusters—such as AI applications, technology industry trends, and consumer electronics discussions—confirming its applicability to real-world social media data.

Suggested Citation

  • Yuanmin Zhang & Hao Li & Chunzhi Xie & Yong Huang & Zhenyi Wu & Yanjun Li, 2026. "Contrastive learning with mutual information enhancement and negative sample augmentation combined with KAN for text clustering," PLOS ONE, Public Library of Science, vol. 21(7), pages 1-35, July.
  • Handle: RePEc:plo:pone00:0354718
    DOI: 10.1371/journal.pone.0354718
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354718
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0354718&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0354718?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0354718. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.