IDEAS home Printed from https://ideas.repec.org/a/sae/somere/v55y2026i2p501-567.html

Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning

Author

Listed:
  • Youngjin Chae
  • Thomas Davidson

Abstract

Large language models (LLMs) have tremendous potential for social science research as they are trained on vast amounts of text and can generalize to many tasks. We explore the use of LLMs for supervised text classification, specifically the application to stance detection, which involves detecting attitudes and opinions in texts. We examine the performance of these models across different architectures, training regimes, and task specifications. We compare 10 models ranging in size from tens of millions to hundreds of billions of parameters and test four distinct training regimes: Prompt-based zero-shot learning and few-shot learning, fine-tuning, and instruction-tuning, which combines prompting and fine-tuning. The largest, most powerful models generally offer the best predictive performance even with little or no training examples, but fine-tuning smaller models is a competitive solution due to their relatively high accuracy and low cost. Instruction-tuning the latest generative LLMs expands the scope of text classification, enabling applications to more complex tasks than previously feasible. We offer practical recommendations on the use of LLMs for text classification in sociological research and discuss their limitations and challenges. Ultimately, LLMs can make text classification and other text analysis methods more accurate, accessible, and adaptable, opening new possibilities for computational social science.

Suggested Citation

  • Youngjin Chae & Thomas Davidson, 2026. "Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning," Sociological Methods & Research, , vol. 55(2), pages 501-567, May.
  • Handle: RePEc:sae:somere:v:55:y:2026:i:2:p:501-567
    DOI: 10.1177/00491241251325243
    as

    Download full text from publisher

    File URL: https://journals.sagepub.com/doi/10.1177/00491241251325243
    Download Restriction: no

    File URL: https://libkey.io/10.1177/00491241251325243?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Bestvater, Samuel E. & Monroe, Burt L., 2023. "Sentiment is Not Stance: Target-Aware Opinion Classification for Political Text Analysis," Political Analysis, Cambridge University Press, vol. 31(2), pages 235-256, April.
    2. Miller, Blake & Linder, Fridolin & Mebane, Walter R., 2020. "Active Learning Approaches for Labeling Text: Review and Assessment of the Performance of Active Learning Approaches," Political Analysis, Cambridge University Press, vol. 28(4), pages 532-551, October.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Sandra Wankmüller, 2023. "A comparison of approaches for imbalanced classification problems in the context of retrieving relevant documents for an analysis," Journal of Computational Social Science, Springer, vol. 6(1), pages 91-163, April.
    2. Karina Shyrokykh & Max Girnyk & Lisa Dellmuth, 2023. "Short text classification with machine learning in the social sciences: The case of climate change on Twitter," PLOS ONE, Public Library of Science, vol. 18(9), pages 1-26, September.
    3. Yang, Yiwen & Lin, Yi-Wei & Cheng, Li-Chen, 2025. "Impact of real-time public sentiment on herding behavior in Taiwan's stock market: Insights across investor types and industries," International Review of Economics & Finance, Elsevier, vol. 102(C).
    4. Sandra Wankmüller, 2024. "Introduction to Neural Transfer Learning With Transformers for Social Science Text Analysis," Sociological Methods & Research, , vol. 53(4), pages 1676-1752, November.
    5. Gründler, Klaus & Potrafke, Niklas & Wochner, Timo, 2025. "Outside employment and parliamentary priorities," Journal of Economic Behavior & Organization, Elsevier, vol. 237(C).
    6. Chendi Wang & Argyrios Altiparmakis, 2025. "What happened to Putin’s friends? The radical right’s reaction to the Russian invasion on social media," European Union Politics, , vol. 26(2), pages 393-417, June.
    7. Karimi Motahhar, Vahid & Gruca, Thomas S. & Tavakoli, Mohammad Hosein, 2025. "Emotions and the status quo: The anti-incumbency bias in political prediction markets," International Journal of Forecasting, Elsevier, vol. 41(2), pages 571-579.
    8. Jonathan Adkins & Ali Al Bataineh & Majd Khalaf, 2024. "Identifying Persons of Interest in Digital Forensics Using NLP-Based AI," Future Internet, MDPI, vol. 16(11), pages 1-19, November.
    9. Leek, Lauren Caroline & Bischl, Simeon, 2024. "How Central Bank Independence Shapes Monetary Policy Communication: A Large Language Model Application," SocArXiv yrhka, Center for Open Science.
    10. Leek, Lauren & Bischl, Simeon, 2025. "How central bank independence shapes monetary policy communication: A Large Language Model application," European Journal of Political Economy, Elsevier, vol. 87(C).
    11. repec:osf:socarx:yrhka_v1 is not listed on IDEAS
    12. R. Michael Alvarez & Jacob Morrier, 2024. "Measuring the Quality of Answers in Political Q&As with Large Language Models," Papers 2404.08816, arXiv.org, revised Feb 2025.
    13. Luca Bellodi & Massimo Morelli & Jorg L. Spenkuch & Edoardo Teso & Matia Vannoni & Guo Xu, 2026. "Personnel is Policy: Delegation and Political Misalignment in the Rulemaking Process," NBER Working Papers 34932, National Bureau of Economic Research, Inc.
    14. Tuan Duy Nguyen & Duc Minh Nguyen & Huu Manh Nguyen & Thi Quynh Giang Nguyen, 2025. "Topic classification of vietnamese product reviews in e-commerce using PhoBERT," Journal of Marketing Analytics, Palgrave Macmillan, vol. 13(2), pages 371-385, June.
    15. Stephanie Ng & James Zhang & Samson Yu & Asim Bhatti & Kathryn Backholer & C. P. Lim, 2025. "Stance classification: a comparative study and use case on Australian parliamentary debates," Journal of Computational Social Science, Springer, vol. 8(2), pages 1-37, May.
    16. Paweł Matuszewski, 2023. "How to prepare data for the automatic classification of politically related beliefs expressed on Twitter? The consequences of researchers’ decisions on the number of coders, the algorithm learning procedure, and the pre-processing steps on the perfor," Quality & Quantity: International Journal of Methodology, Springer, vol. 57(1), pages 301-321, February.

    More about this item

    Keywords

    ;
    ;
    ;
    ;
    ;

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:sae:somere:v:55:y:2026:i:2:p:501-567. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: SAGE Publications (email available below). General contact details of provider: .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.