IDEAS home Printed from https://ideas.repec.org/a/bjf/journl/v11y2026i5p309-323.html

HPCM: A Hybrid Multi-Layered Machine Learning Pipeline for Plagiarism Content Matching with Dynamic Threshold Calibration

Author

Listed:
  • Piyush Chavan

    (Computer Science, Pune, Maharashtra, India)

  • Prof.Moushmee Kuri

    (Computer Science, Pune, Maharashtra, India)

  • Tanvi Bokade

    (Computer Science, Pune, Maharashtra, India)

  • Pushkar Thombare

    (Computer Science, Pune, Maharashtra, India)

Abstract

Conventional approaches to detecting plagiarism involve mainly string-matching and n-gram fingerprinting methods, which can detect plagiarised documents involving verbatim plagiarism, but they cannot catch paraphrasing, synonym substitutions, or imitations of writing styles. Such shortcomings have now gained importance due to developments of sophisticated intelligent paraphrasing and the use of advanced large language models, which help evade detection by conventional approaches. In this research, we present HPCM, an end-to-end plagiarism detection system that utilises a nine-module machine-learning-based pipeline combining three analysis components: the first is the cosine similarity of terms using the TF-IDF method, secondly, embedding-based semantic similarity using the all-miniLM-L6-v2 model, and thirdly, stylistic similarity based on the analysis of POS Distribution, Type Token Ratio, and Sentence length statistics. These results are combined through the application of a weighted sum fusion function that gives greater emphasis to the semantic similarity score. Additionally, a novel Dynamic Similarity Calibration (DSC) module adjusts the plagiarism score per pair based on the relative length of documents, their vocabulary richness, and topic similarity. Experiments conducted over four different categories of plagiarism reveal that HPCM scores 69.0% in detecting paraphrases compared to 24.9% by conventional approaches, showing a remarkable 44.1 percentage point improvement. It is implemented as a microservices system on Vercel, Render, Hugging Face Spaces, and MongoDB Atlas, proving the practicality of using multilayered neural models for detecting plagiarism even with only free-tier cloud resources. The source code, along with the testing data, is publicly available.

Suggested Citation

  • Piyush Chavan & Prof.Moushmee Kuri & Tanvi Bokade & Pushkar Thombare, 2026. "HPCM: A Hybrid Multi-Layered Machine Learning Pipeline for Plagiarism Content Matching with Dynamic Threshold Calibration," International Journal of Research and Innovation in Applied Science, International Journal of Research and Innovation in Applied Science (IJRIAS), vol. 11(5), pages 309-323, May.
  • Handle: RePEc:bjf:journl:v:11:y:2026:i:5:p:309-323
    as

    Download full text from publisher

    File URL: https://rsisinternational.org/journals/ijrias/uploads/vol11-iss5-pg309-323-202605_pdf.pdf
    Download Restriction: no

    File URL: https://rsisinternational.org/journals/ijrias/view/hpcm-a-hybrid-multi-layered-machine-learning-pipeline-for-plagiarism-content-matching-with-dynamic-threshold-calibration/
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Efstathios Stamatatos, 2009. "A survey of modern authorship attribution methods," Journal of the American Society for Information Science and Technology, Association for Information Science & Technology, vol. 60(3), pages 538-556, March.
    2. Scott Deerwester & Susan T. Dumais & George W. Furnas & Thomas K. Landauer & Richard Harshman, 1990. "Indexing by latent semantic analysis," Journal of the American Society for Information Science, Association for Information Science & Technology, vol. 41(6), pages 391-407, September.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Kargin, Vladislav, 2016. "On variation of word frequencies in Russian literary texts," Physica A: Statistical Mechanics and its Applications, Elsevier, vol. 445(C), pages 328-334.
    2. Irina Wedel & Michael Palk & Stefan Voß, 2022. "A Bilingual Comparison of Sentiment and Topics for a Product Event on Twitter," Information Systems Frontiers, Springer, vol. 24(5), pages 1635-1646, October.
    3. Ca' Zorzi, Michele & Manu, Ana-Simona & Lopardo, Gianluigi, 2025. "Verba volant, transcripta manent: what corporate earnings calls reveal about the AI stock rally," Working Paper Series 3093, European Central Bank.
    4. Mohammed Salem Binwahlan, 2023. "Polynomial Networks Model for Arabic Text Summarization," International Journal of Research and Scientific Innovation, International Journal of Research and Scientific Innovation (IJRSI), vol. 10(2), pages 74-84, February.
    5. Curci, Ylenia & Mongeau Ospina, Christian A., 2016. "Investigating biofuels through network analysis," Energy Policy, Elsevier, vol. 97(C), pages 60-72.
    6. Chao Wei & Senlin Luo & Xincheng Ma & Hao Ren & Ji Zhang & Limin Pan, 2016. "Locally Embedding Autoencoders: A Semi-Supervised Manifold Learning Approach of Document Representation," PLOS ONE, Public Library of Science, vol. 11(1), pages 1-20, January.
    7. Marcello D’Amato & Francesco Flaviano Russo, 2026. "Cultural doorways in the barriers to development," Journal of Economic Growth, Springer, vol. 31(1), pages 125-178, March.
    8. Pietro Fera & Nicola Moscariello & Gianmarco Salzillo & Emilio Farina, 2025. "Towards the Regulation of Non‐Financial Reporting: The Impact on Environmental Disclosure Within the Oil and Gas Sector," Corporate Social Responsibility and Environmental Management, John Wiley & Sons, vol. 32(3), pages 4053-4067, May.
    9. Maksym Polyakov & Morteza Chalak & Md. Sayed Iftekhar & Ram Pandit & Sorada Tapsuwan & Fan Zhang & Chunbo Ma, 2018. "Authorship, Collaboration, Topics, and Research Gaps in Environmental and Resource Economics 1991–2015," Environmental & Resource Economics, Springer;European Association of Environmental and Resource Economists, vol. 71(1), pages 217-239, September.
    10. Ding, Ying, 2011. "Community detection: Topological vs. topical," Journal of Informetrics, Elsevier, vol. 5(4), pages 498-514.
    11. Klaus Gugler & Florian Szücs & Ulrich Wohak, 2023. "Start-up Acquisitions, Venture Capital and Innovation: A Comparative Study of Google, Apple, Facebook, Amazon and Microsoft," Department of Economics Working Papers wuwp340, Vienna University of Economics and Business, Department of Economics.
    12. Md Nazrul Islam & Md Mofazzal Hossain & Md Shafayet Shahed Ornob, 2024. "Business research on Industry 4.0: a systematic review using topic modelling approach," Future Business Journal, Springer, vol. 10(1), pages 1-15, December.
    13. Juan Shi & Kin Keung Lai & Ping Hu & Gang Chen, 2018. "Factors dominating individual information disseminating behavior on social networking sites," Information Technology and Management, Springer, vol. 19(2), pages 121-139, June.
    14. Ganesh Dash & Chetan Sharma & Shamneesh Sharma, 2023. "Sustainable Marketing and the Role of Social Media: An Experimental Study Using Natural Language Processing (NLP)," Sustainability, MDPI, vol. 15(6), pages 1-16, March.
    15. repec:osf:socarx:49qxk_v1 is not listed on IDEAS
    16. Paola Cerchiello & Giancarlo Nicola, 2018. "Assessing News Contagion in Finance," Econometrics, MDPI, vol. 6(1), pages 1-19, February.
    17. Shr-Wei Kao & Pin Luarn, 2020. "Topic Modeling Analysis of Social Enterprises: Twitter Evidence," Sustainability, MDPI, vol. 12(8), pages 1-20, April.
    18. repec:plo:pone00:0197933 is not listed on IDEAS
    19. Gissler, Stefan & Oldfather, Jeremy & Ruffino, Doriana, 2016. "Lending on hold: Regulatory uncertainty and bank lending standards," Journal of Monetary Economics, Elsevier, vol. 81(C), pages 89-101.
    20. Nils-Axel M?rner, 2018. "Evaluation of the Performance and Efficiency of the Automated Linguistic Features for Author Identification in Short Text Messages Using Different Variable Selection Techniques," Studies in Media and Communication, Redfame publishing, vol. 6(2), pages 83-102, December.
    21. Wittek, Peter, 2013. "Two-way incremental seriation in the temporal domain with three-dimensional visualization: Making sense of evolving high-dimensional datasets," Computational Statistics & Data Analysis, Elsevier, vol. 66(C), pages 193-201.
    22. Alina Evstigneeva & Mark Sidorovskiy, 2021. "Assessment of Clarity of Bank of Russia Monetary Policy Communication by Neural Network Approach," Russian Journal of Money and Finance, Bank of Russia, vol. 80(3), pages 3-33, September.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bjf:journl:v:11:y:2026:i:5:p:309-323. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Dr. Renu Malsaria (email available below). General contact details of provider: https://rsisinternational.org/journals/ijrias/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.