IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2103.11933.html
   My bibliography  Save this paper

PatentSBERTa: A Deep NLP based Hybrid Model for Patent Distance and Classification using Augmented SBERT

Author

Listed:
  • Hamid Bekamiri
  • Daniel S. Hain
  • Roman Jurowetzki

Abstract

This study provides an efficient approach for using text data to calculate patent-to-patent (p2p) technological similarity, and presents a hybrid framework for leveraging the resulting p2p similarity for applications such as semantic search and automated patent classification. We create embeddings using Sentence-BERT (SBERT) based on patent claims. We leverage SBERTs efficiency in creating embedding distance measures to map p2p similarity in large sets of patent data. We deploy our framework for classification with a simple Nearest Neighbors (KNN) model that predicts Cooperative Patent Classification (CPC) of a patent based on the class assignment of the K patents with the highest p2p similarity. We thereby validate that the p2p similarity captures their technological features in terms of CPC overlap, and at the same demonstrate the usefulness of this approach for automatic patent classification based on text data. Furthermore, the presented classification framework is simple and the results easy to interpret and evaluate by end-users. In the out-of-sample model validation, we are able to perform a multi-label prediction of all assigned CPC classes on the subclass (663) level on 1,492,294 patents with an accuracy of 54% and F1 score > 66%, which suggests that our model outperforms the current state-of-the-art in text-based multi-label and multi-class patent classification. We furthermore discuss the applicability of the presented framework for semantic IP search, patent landscaping, and technology intelligence. We finally point towards a future research agenda for leveraging multi-source patent embeddings, their appropriateness across applications, as well as to improve and validate patent embeddings by creating domain-expert curated Semantic Textual Similarity (STS) benchmark datasets.

Suggested Citation

  • Hamid Bekamiri & Daniel S. Hain & Roman Jurowetzki, 2021. "PatentSBERTa: A Deep NLP based Hybrid Model for Patent Distance and Classification using Augmented SBERT," Papers 2103.11933, arXiv.org, revised Oct 2021.
  • Handle: RePEc:arx:papers:2103.11933
    as

    Download full text from publisher

    File URL: http://arxiv.org/pdf/2103.11933
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Grigorios Tsoumakas & Ioannis Katakis, 2007. "Multi-Label Classification: An Overview," International Journal of Data Warehousing and Mining (IJDWM), IGI Global, vol. 3(3), pages 1-13, July.
    2. Shaobo Li & Jie Hu & Yuxin Cui & Jianjun Hu, 2018. "DeepPatent: patent classification with convolutional neural networks and word embedding," Scientometrics, Springer;Akadémiai Kiadó, vol. 117(2), pages 721-744, November.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Kang, Byeongwoo & Bekkers, Rudi, 2022. "The determinants of parallel invention : Measuring the role of information sharing and personal interaction between inventors," IIR Working Paper 22-06, Institute of Innovation Research, Hitotsubashi University.
    2. Jeon, Eunji & Yoon, Naeun & Sohn, So Young, 2023. "Exploring new digital therapeutics technologies for psychiatric disorders using BERTopic and PatentSBERTa," Technological Forecasting and Social Change, Elsevier, vol. 186(PA).

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Anqi Ma & Yu Liu & Xiujuan Xu & Tao Dong, 2021. "A deep-learning based citation count prediction model with paper metadata semantic features," Scientometrics, Springer;Akadémiai Kiadó, vol. 126(8), pages 6803-6823, August.
    2. Radu Cristian Alexandru Iacob & Vlad Cristian Monea & Dan Rădulescu & Andrei-Florin Ceapă & Traian Rebedea & Ștefan Trăușan-Matu, 2020. "AlgoLabel: A Large Dataset for Multi-Label Classification of Algorithmic Challenges," Mathematics, MDPI, vol. 8(11), pages 1-18, November.
    3. Azzini, Antonia & Cortesi, Nicola & Marrara, Stefania & Topalović, Amir, 2019. "A Multi-Label Machine Learning Approach to Support Pathologist's Histological Analysis," Proceedings of the ENTRENOVA - ENTerprise REsearch InNOVAtion Conference (2019), Rovinj, Croatia, in: Proceedings of the ENTRENOVA - ENTerprise REsearch InNOVAtion Conference, Rovinj, Croatia, 12-14 September 2019, pages 197-208, IRENET - Society for Advancing Innovation and Research in Economy, Zagreb.
    4. Xueying Zhang & Qinbao Song, 2015. "A Multi-Label Learning Based Kernel Automatic Recommendation Method for Support Vector Machine," PLOS ONE, Public Library of Science, vol. 10(4), pages 1-30, April.
    5. Junming Yin & Jerry Luo & Susan A. Brown, 2021. "Learning from Crowdsourced Multi-labeling: A Variational Bayesian Approach," Information Systems Research, INFORMS, vol. 32(3), pages 752-773, September.
    6. Jeon, Eunji & Yoon, Naeun & Sohn, So Young, 2023. "Exploring new digital therapeutics technologies for psychiatric disorders using BERTopic and PatentSBERTa," Technological Forecasting and Social Change, Elsevier, vol. 186(PA).
    7. Arousha Haghighian Roudsari & Jafar Afshar & Wookey Lee & Suan Lee, 2022. "PatentNet: multi-label classification of patent documents using deep learning based language understanding," Scientometrics, Springer;Akadémiai Kiadó, vol. 127(1), pages 207-231, January.
    8. Yuan Zhou & Fang Dong & Yufei Liu & Liang Ran, 2021. "A deep learning framework to early identify emerging technologies in large-scale outlier patents: an empirical study of CNC machine tool," Scientometrics, Springer;Akadémiai Kiadó, vol. 126(2), pages 969-994, February.
    9. Choi, Seokkyu & Lee, Hyeonju & Park, Eunjeong & Choi, Sungchul, 2022. "Deep learning for patent landscaping using transformer and graph embedding," Technological Forecasting and Social Change, Elsevier, vol. 175(C).
    10. Peng Shao & Runhua Tan & Qingjin Peng & Wendan Yang & Fang Liu, 2023. "An Integrated Method to Acquire Technological Evolution Potential to Stimulate Innovative Product Design," Mathematics, MDPI, vol. 11(3), pages 1-24, January.
    11. Ascione, Grazia Sveva, 2023. "Technological diversity to address complex challenges: the contribution of American universities to sdgs," MPRA Paper 119452, University Library of Munich, Germany.
    12. Tadeusz A. Grzeszczyk & Michal K. Grzeszczyk, 2021. "Improving the Discovery of Technological Opportunities Using Patent Classification Based on Explainable Neural Networks," European Research Studies Journal, European Research Studies Journal, vol. 0(3), pages 402-409.
    13. Liang Chen & Shuo Xu & Lijun Zhu & Jing Zhang & Xiaoping Lei & Guancan Yang, 2020. "A deep learning based method for extracting semantic information from patent documents," Scientometrics, Springer;Akadémiai Kiadó, vol. 125(1), pages 289-312, October.
    14. Bocheng Li & Yunqiu Zhang & Xusheng Wu, 2022. "DLKN-MLC: A Disease Prediction Model via Multi-Label Learning," IJERPH, MDPI, vol. 19(15), pages 1-15, August.
    15. Puccetti, Giovanni & Giordano, Vito & Spada, Irene & Chiarello, Filippo & Fantoni, Gualtiero, 2023. "Technology identification from patent texts: A novel named entity recognition method," Technological Forecasting and Social Change, Elsevier, vol. 186(PB).
    16. Chaker Jebari, 2016. "Multi-Label Genre Classification of Web Pages Using an Adaptive Centroid-Based Classifier," Journal of Information & Knowledge Management (JIKM), World Scientific Publishing Co. Pte. Ltd., vol. 15(01), pages 1-21, March.
    17. Francisco J. Ribadas-Pena & Shuyuan Cao & Víctor M. Darriba Bilbao, 2022. "Improving Large-Scale k -Nearest Neighbor Text Categorization with Label Autoencoders," Mathematics, MDPI, vol. 10(16), pages 1-22, August.
    18. Mark Bukowski & Sandra Geisler & Thomas Schmitz-Rode & Robert Farkas, 2020. "Feasibility of activity-based expert profiling using text mining of scientific publications and patents," Scientometrics, Springer;Akadémiai Kiadó, vol. 123(2), pages 579-620, May.
    19. Tao Shu & Zhiyi Wang & Huading Jia & Wenjin Zhao & Jixian Zhou & Tao Peng, 2022. "Consumers’ Opinions towards Public Health Effects of Online Games: An Empirical Study Based on Social Media Comments in China," IJERPH, MDPI, vol. 19(19), pages 1-19, October.
    20. Meindl, Benjamin & Ayala, Néstor Fabián & Mendonça, Joana & Frank, Alejandro G., 2021. "The four smarts of Industry 4.0: Evolution of ten years of research and future perspectives," Technological Forecasting and Social Change, Elsevier, vol. 168(C).

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2103.11933. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: http://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.