IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0355603.html

A novel hybrid model for identifying the most informative instances for improving text data classification

Author

Listed:
  • Amira Abdelwahab
  • Mohamed Salama

Abstract

The rapid growth of user-generated textual content on the internet has intensified the need for accurate and scalable text classification methods. However, supervised learning approaches remain heavily constrained by the high cost and effort required for manual data annotation, particularly in large and heterogeneous datasets. To address this challenge, this paper proposes a novel hybrid active learning framework for efficient classification of unlabeled text data. The proposed approach integrates multiple classical machine learning classifiers—Support Vector Machines, Logistic Regression, Naive Bayes, and Random Forest—within a hybrid ensemble architecture, combined with a pool-based active learning strategy to iteratively select the most informative unlabeled instances for annotation. Textual data are transformed into numerical representations using several feature extraction techniques, including Bag-of-Words, TF-IDF, Word2Vec, and BERT-based embeddings, allowing for a comprehensive evaluation of representation effectiveness. Extensive experiments are conducted on four diverse benchmark datasets from healthcare, finance, spam detection, and e-commerce domains. The results consistently demonstrate that the proposed hybrid active learning model outperforms traditional ensemble classifiers across all datasets and evaluation metrics. In particular, TF-IDF-based hybrid ensembles achieve the highest gains in accuracy, precision, recall, and F1 score, while requiring substantially fewer labeled instances. Furthermore, the proposed framework exhibits strong robustness in imbalanced classification scenarios, significantly improving minority class detection. Overall, the findings confirm that combining hybrid ensemble learning with active learning offers an effective, lightweight, and cost-efficient alternative to purely transformer-based approaches, making it well-suited for real-world text classification tasks where labeled data are scarce or expensive.

Suggested Citation

  • Amira Abdelwahab & Mohamed Salama, 2026. "A novel hybrid model for identifying the most informative instances for improving text data classification," PLOS ONE, Public Library of Science, vol. 21(8), pages 1-20, August.
  • Handle: RePEc:plo:pone00:0355603
    DOI: 10.1371/journal.pone.0355603
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0355603
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0355603&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0355603?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0355603. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.