IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0267590.html

Detection of changes in literary writing style using N-grams as style markers and supervised machine learning

Author

Listed:
  • Germán Ríos-Toledo
  • Juan Pablo Francisco Posadas-Durán
  • Grigori Sidorov
  • Noé Alejandro Castro-Sánchez

Abstract

The analysis of an author’s writing style implies the characterization and identification of the style in terms of a set of features commonly called linguistic features. The analysis can be extrinsic, where the style of an author can be compared with other authors, or intrinsic, where the style of an author is identified through different stages of his life. Intrinsic analysis has been used, for example, to detect mental illness and the effects of aging. A key element of the analysis is the style markers used to model the author’s writing patterns. The style markers should handle diachronic changes and be thematic independent. One of the most commonly used style marker in extrinsic style analysis is n-gram. In this paper, we present the evaluation of traditional n-grams (words and characters) and dependency tree syntactic n-grams to solve the task of detecting changes in writing style over time. Our corpus consisted of novels by eleven English-speaking authors. The novels of each author were organized chronologically from the oldest to the most recent work according to the date of publication. Subsequently, two stages were defined: initial and final. In each stage three novels were assigned, novels of the initial stage corresponded to the oldest and those at the final stage to the most recent novels. To analyze changes in the writing style, novels were characterized by using four types of n-grams: characters, words, Part-Of-Speech (POS) tags and syntactic relations n-grams. Experiments were performed with a Logistic Regression classifier. Dimension reduction techniques such as Principal Component Analysis (PCA) and Latent Semantic Analysis (LSA) algorithms were evaluated. The results obtained with the different n-grams indicated that all authors presented significant changes in writing style over time. In addition, representations using n-grams of syntactic relations have achieved competitive results among different authors.

Suggested Citation

  • Germán Ríos-Toledo & Juan Pablo Francisco Posadas-Durán & Grigori Sidorov & Noé Alejandro Castro-Sánchez, 2022. "Detection of changes in literary writing style using N-grams as style markers and supervised machine learning," PLOS ONE, Public Library of Science, vol. 17(7), pages 1-24, July.
  • Handle: RePEc:plo:pone00:0267590
    DOI: 10.1371/journal.pone.0267590
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0267590
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0267590&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0267590?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    References listed on IDEAS

    as
    1. Matthias Schonlau & Nick Guenther & Ilia Sucholutsky, 2017. "Text mining with n-gram variables," Stata Journal, StataCorp LLC, vol. 17(4), pages 866-881, December.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Cristina Cattaneo & Daniela Grieco & Nicola Lacetera & Mario Macis, 2025. "Out‐group penalties in refugee assistance: a survey experiment," Scandinavian Journal of Economics, Wiley Blackwell, vol. 127(4), pages 697-741, October.
    2. Paul M. Anglin & Yanmin Gao, 2023. "Value of Communication and Social Media: An Equilibrium Theory of Messaging," The Journal of Real Estate Finance and Economics, Springer, vol. 66(4), pages 861-903, May.
    3. Brown, Martin & Schmitz, Jan & Zehnder, Christian, 2024. "Communication and hidden action: A credit market experiment," Journal of Economic Behavior & Organization, Elsevier, vol. 218(C), pages 423-455.
    4. Fatma Yiğit Açikgöz & Mehmet Kayakuş & Georgiana Moiceanu & Nesrin Sönmez, 2024. "A New Approach to Assess Sustainable Corporate Reputation with Citizen Comments Using Machine Learning and Natural Language Processing," Sustainability, MDPI, vol. 16(22), pages 1-19, November.
    5. Christopher Haynes & Marco A. Palomino & Liz Stuart & David Viira & Frances Hannon & Gemma Crossingham & Kate Tantam, 2022. "Automatic Classification of National Health Service Feedback," Mathematics, MDPI, vol. 10(6), pages 1-23, March.
    6. Muhammed-Fatih Kaya, 2022. "Pattern Labelling of Business Communication Data," Group Decision and Negotiation, Springer, vol. 31(6), pages 1203-1234, December.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0267590. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.