The (real) need for a human touch: testing a human–machine hybrid topic classification workflow on a New York Times corpus

My bibliography Save this article

The (real) need for a human touch: testing a human–machine hybrid topic classification workflow on a New York Times corpus

Author

Listed:

Miklos Sebők
(Centre for Social Sciences)
Zoltán Kacsuk
(Centre for Social Sciences
Hochschule der Medien)
Ákos Máté
(Centre for Social Sciences)

Registered:

Abstract

The classification of the items of ever-increasing textual databases has become an important goal for a number of research groups active in the field of computational social science. Due to the increased amount of text data there is a growing number of use-cases where the initial effort of human classifiers was successfully augmented using supervised machine learning (SML). In this paper, we investigate such a hybrid workflow solution classifying the lead paragraphs of New York Times front-page articles from 1996 to 2006 according to policy topic categories (such as education or defense) of the Comparative Agendas Project (CAP). The SML classification is conducted in multiple rounds and, within each round, we run the SML algorithm on n samples and n times if the given algorithm is non-deterministic (e.g., SVM). If all the SML predictions point towards a single label for a document, then it is classified as such (this approach is also called a “voting ensemble"). In the second step, we explore several scenarios, ranging from using the SML ensemble without human validation to incorporating active learning. Using these scenarios, we can quantify the gains from the various workflow versions. We find that using human coding and validation combined with an ensemble SML hybrid approach can reduce the need for human coding while maintaining very high precision rates and offering a modest to a good level of recall. The modularity of this hybrid workflow allows for various setups to address the idiosyncratic resource bottlenecks that a large-scale text classification project might face.

Suggested Citation

Miklos Sebők & Zoltán Kacsuk & Ákos Máté, 2022. "The (real) need for a human touch: testing a human–machine hybrid topic classification workflow on a New York Times corpus," Quality & Quantity: International Journal of Methodology, Springer, vol. 56(5), pages 3621-3643, October.

Handle: RePEc:spr:qualqt:v:56:y:2022:i:5:d:10.1007_s11135-021-01287-4
DOI: 10.1007/s11135-021-01287-4

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

Sebők, Miklós & Kacsuk, Zoltán, 2021. "The Multiclass Classification of Newspaper Articles with Machine Learning: The Hybrid Binary Snowball Approach," Political Analysis, Cambridge University Press, vol. 29(2), pages 236-249, April.
Stuart N. Soroka & Dominik A. Stecula & Christopher Wlezien, 2015. "It's (Change in) the (Future) Economy, Stupid: Economic Indicators, the Media, and Public Opinion," American Journal of Political Science, John Wiley & Sons, vol. 59(2), pages 457-474, February.
Grimmer, Justin & Stewart, Brandon M., 2013. "Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts," Political Analysis, Cambridge University Press, vol. 21(3), pages 267-297, July.
Denny, Matthew J. & Spirling, Arthur, 2018. "Text Preprocessing For Unsupervised Learning: Why It Matters, When It Misleads, And What To Do About It," Political Analysis, Cambridge University Press, vol. 26(2), pages 168-189, April.
Mike Thelwall & Kevan Buckley & Georgios Paltoglou, 2012. "Sentiment strength detection for the social web," Journal of the Association for Information Science & Technology, Association for Information Science & Technology, vol. 63(1), pages 163-173, January.
Lucas, Christopher & Nielsen, Richard A. & Roberts, Margaret E. & Stewart, Brandon M. & Storer, Alex & Tingley, Dustin, 2015. "Computer-Assisted Text Analysis for Comparative Politics," Political Analysis, Cambridge University Press, vol. 23(2), pages 254-277, April.
Adam Bonica, 2018. "Inferring Roll‐Call Scores from Campaign Contributions Using Supervised Machine Learning," American Journal of Political Science, John Wiley & Sons, vol. 62(4), pages 830-848, October.
Peterson, Andrew & Spirling, Arthur, 2018. "Classification Accuracy as a Substantive Quantity of Interest: Measuring Polarization in Westminster Systems," Political Analysis, Cambridge University Press, vol. 26(1), pages 120-128, January.
Mike Thelwall & Kevan Buckley & Georgios Paltoglou, 2012. "Sentiment strength detection for the social web," Journal of the American Society for Information Science and Technology, Association for Information Science & Technology, vol. 63(1), pages 163-173, January.

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Mohamed M. Mostafa, 2023. "A one-hundred-year structural topic modeling analysis of the knowledge structure of international management research," Quality & Quantity: International Journal of Methodology, Springer, vol. 57(4), pages 3905-3935, August.
Seraphine F. Maerz & Carsten Q. Schneider, 2020. "Comparing public communication in democracies and autocracies: automated text analyses of speeches by heads of government," Quality & Quantity: International Journal of Methodology, Springer, vol. 54(2), pages 517-545, April.
Bastiaan Bruinsma & Moa Johansson, 2024. "Finding the structure of parliamentary motions in the Swedish Riksdag 1971–2015," Quality & Quantity: International Journal of Methodology, Springer, vol. 58(4), pages 3275-3301, August.
Jaehwan LIM & Asei ITO & Hongyong ZHANG, 2023. "Policy Agenda and Trajectory of the Xi Jinping Administration: Textual Evidence from 2012 to 2022," Policy Discussion Papers 23008, Research Institute of Economy, Trade and Industry (RIETI).
Karina Shyrokykh & Max Girnyk & Lisa Dellmuth, 2023. "Short text classification with machine learning in the social sciences: The case of climate change on Twitter," PLOS ONE, Public Library of Science, vol. 18(9), pages 1-26, September.
Martin Haselmayer & Marcelo Jenny, 2017. "Sentiment analysis of political communication: combining a dictionary approach with crowdcoding," Quality & Quantity: International Journal of Methodology, Springer, vol. 51(6), pages 2623-2646, November.
repec:osf:socarx:yfzsh_v1 is not listed on IDEAS
Karell, Daniel & Freedman, Michael Raphael, 2019. "Rhetorics of Radicalism," SocArXiv yfzsh, Center for Open Science.
Michal OvÃ¡dek & Nicolas Lampach & Arthur Dyevre, 2020. "Whatâ€™s the talk in Brussels? Leveraging daily news coverage to measure issue attention in the European Union," European Union Politics, , vol. 21(2), pages 204-232, June.
Purwoko Haryadi Santoso & Edi Istiyono & Haryanto & Wahyu Hidayatulloh, 2022. "Thematic Analysis of Indonesian Physics Education Research Literature Using Machine Learning," Data, MDPI, vol. 7(11), pages 1-41, October.
Agrawal, Shiv Ratan & Mittal, Divya, 2022. "Optimizing customer engagement content strategy in retail and E-tail: Available on online product review videos," Journal of Retailing and Consumer Services, Elsevier, vol. 67(C).
Ferrara, Federico M. & Masciandaro, Donato & Moschella, Manuela & Romelli, Davide, 2022. "Political voice on monetary policy: Evidence from the parliamentary hearings of the European Central Bank," European Journal of Political Economy, Elsevier, vol. 74(C).
- Federico M. Ferrara & Donato Masciandaro & Manuela Moschella & Davide Romelli, 2021. "Political Voice on Monetary Policy: Evidence from the Parliamentary Hearings of the European Central Bank," BAFFI CAREFIN Working Papers 21159, BAFFI CAREFIN, Centre for Applied Research on International Markets Banking Finance and Regulation, Universita' Bocconi, Milano, Italy.
- Ferrara, Federico M. & Masciandaro, Donato & Moschella, Manuela & Romelli, Davide, 2022. "Political voice on monetary policy: evidence from the parliamentary hearings of the European Central Bank," LSE Research Online Documents on Economics 114278, London School of Economics and Political Science, LSE Library.
Camilla Salvatore & Silvia Biffignandi & Annamaria Bianchi, 2022. "Corporate Social Responsibility Activities Through Twitter: From Topic Model Analysis to Indexes Measuring Communication Characteristics," Social Indicators Research: An International and Interdisciplinary Journal for Quality-of-Life Measurement, Springer, vol. 164(3), pages 1217-1248, December.
Jason Anastasopoulos & George J. Borjas & Gavin G. Cook & Michael Lachanski, 2018. "Job Vacancies, the Beveridge Curve, and Supply Shocks: The Frequency and Content of Help-Wanted Ads in Pre- and Post-Mariel Miami," NBER Working Papers 24580, National Bureau of Economic Research, Inc.
- Anastasopoulos, Jason & Borjas, George J. & Cook, Gavin G. & Lachanski, Michael, 2019. "Job Vacancies, the Beveridge Curve, and Supply Shocks: The Frequency and Content of Help-Wanted Ads in Pre- and Post-Mariel Miami," IZA Discussion Papers 12581, Institute of Labor Economics (IZA).
Young Bin Kim & Sang Hyeok Lee & Shin Jin Kang & Myung Jin Choi & Jung Lee & Chang Hun Kim, 2015. "Virtual World Currency Value Fluctuation Prediction System Based on User Sentiment Analysis," PLOS ONE, Public Library of Science, vol. 10(8), pages 1-18, August.
Ping-Yu Hsu & Hong-Tsuen Lei & Shih-Hsiang Huang & Teng Hao Liao & Yao-Chung Lo & Chin-Chun Lo, 2019. "Effects of sentiment on recommendations in social network," Electronic Markets, Springer;IIM University of St. Gallen, vol. 29(2), pages 253-262, June.
Dehler-Holland, Joris & Okoh, Marvin & Keles, Dogan, 2022. "Assessing technology legitimacy with topic models and sentiment analysis – The case of wind power in Germany," Technological Forecasting and Social Change, Elsevier, vol. 175(C).
- Dehler-Holland, Joris & Okoh, Marvin & Keles, Dogan, 2021. "The legitimacy of wind power in Germany," Working Paper Series in Production and Energy 54, Karlsruhe Institute of Technology (KIT), Institute for Industrial Production (IIP).
Luis J. Callarisa-Fiol & Miguel Ángel Moliner-Tena & Rosa Rodríguez-Artola & Javier Sánchez-García, 2023. "Entrepreneurship innovation using social robots in tourism: a social listening study," Review of Managerial Science, Springer, vol. 17(8), pages 2945-2971, November.
Ulrich Fritsche & Johannes Puckelwald, 2018. "Deciphering Professional Forecasters’ Stories - Analyzing a Corpus of Textual Predictions for the German Economy," Macroeconomics and Finance Series 201804, University of Hamburg, Department of Socioeconomics.
Anustubh Agnihotri & Rahul Verma, 2019. "Content Analysis of Digital Text and Its Applications," Studies in Indian Politics, , vol. 7(1), pages 83-89, June.
repec:osf:osfxxx:ghxj8_v1 is not listed on IDEAS
Iasmin Goes, 2023. "Examining the effect of IMF conditionality on natural resource policy," Economics and Politics, Wiley Blackwell, vol. 35(1), pages 227-285, March.

More about this item

Keywords

; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:qualqt:v:56:y:2022:i:5:d:10.1007_s11135-021-01287-4. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

The (real) need for a human touch: testing a human–machine hybrid topic classification workflow on a New York Times corpus

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data