IDEAS home Printed from https://ideas.repec.org/a/wsi/jikmxx/v15y2016i01ns0219649216500088.html
   My bibliography  Save this article

Multi-Label Genre Classification of Web Pages Using an Adaptive Centroid-Based Classifier

Author

Listed:
  • Chaker Jebari

    (IT Department, College of Applied Sciences, IBRI, BOX 516, Sultanate of Oman)

Abstract

This paper proposes an adaptive centroid-based classifier (ACC) for multi-label classification of web pages. Using a set of multi-genre training dataset, ACC constructs a centroid for each genre. To deal with the rapid evolution of web genres, ACC implements an adaptive classification method where web pages are classified one by one. For each web page, ACC calculated its similarity with all genre centroids. Based on this similarity, ACC either adjusts the genre centroid by including the new web page or discards it. A web page is a complex object that contains different sections belonging to different genres. To handle this complexity, ACC implements a multi-label classification where a web page can be assigned to multiple genres at the same time. To improve the performance of genre classification, we propose to aggregate the classifications produced using character n-grams extracted from URL, title, headings and anchors. Experiments conducted using a known multi-label dataset show that ACC outperforms many other multi-label classifiers and has the lowest computational complexity.

Suggested Citation

  • Chaker Jebari, 2016. "Multi-Label Genre Classification of Web Pages Using an Adaptive Centroid-Based Classifier," Journal of Information & Knowledge Management (JIKM), World Scientific Publishing Co. Pte. Ltd., vol. 15(01), pages 1-21, March.
  • Handle: RePEc:wsi:jikmxx:v:15:y:2016:i:01:n:s0219649216500088
    DOI: 10.1142/S0219649216500088
    as

    Download full text from publisher

    File URL: http://www.worldscientific.com/doi/abs/10.1142/S0219649216500088
    Download Restriction: Access to full text is restricted to subscribers

    File URL: https://libkey.io/10.1142/S0219649216500088?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Grigorios Tsoumakas & Ioannis Katakis, 2007. "Multi-Label Classification: An Overview," International Journal of Data Warehousing and Mining (IJDWM), IGI Global, vol. 3(3), pages 1-13, July.
    Full references (including those not matched with items on IDEAS)

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Hanan Al-Mofareji & Mahmoud Kamel & Mohamed Y. Dahab, 2017. "WeDoCWT: A New Method for Web Document Clustering Using Discrete Wavelet Transforms," Journal of Information & Knowledge Management (JIKM), World Scientific Publishing Co. Pte. Ltd., vol. 16(01), pages 1-19, March.
    2. Ruchika Malhotra & Anjali Sharma, 2017. "Quantitative evaluation of web metrics for automatic genre classification of web pages," International Journal of System Assurance Engineering and Management, Springer;The Society for Reliability, Engineering Quality and Operations Management (SREQOM),India, and Division of Operation and Maintenance, Lulea University of Technology, Sweden, vol. 8(2), pages 1567-1579, November.

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Radu Cristian Alexandru Iacob & Vlad Cristian Monea & Dan Rădulescu & Andrei-Florin Ceapă & Traian Rebedea & Ștefan Trăușan-Matu, 2020. "AlgoLabel: A Large Dataset for Multi-Label Classification of Algorithmic Challenges," Mathematics, MDPI, vol. 8(11), pages 1-18, November.
    2. Xueying Zhang & Qinbao Song, 2015. "A Multi-Label Learning Based Kernel Automatic Recommendation Method for Support Vector Machine," PLOS ONE, Public Library of Science, vol. 10(4), pages 1-30, April.
    3. Junming Yin & Jerry Luo & Susan A. Brown, 2021. "Learning from Crowdsourced Multi-labeling: A Variational Bayesian Approach," Information Systems Research, INFORMS, vol. 32(3), pages 752-773, September.
    4. Hamid Bekamiri & Daniel S. Hain & Roman Jurowetzki, 2021. "PatentSBERTa: A Deep NLP based Hybrid Model for Patent Distance and Classification using Augmented SBERT," Papers 2103.11933, arXiv.org, revised Oct 2021.
    5. Yi-Hui Chen & Eric Jui-Lin Lu & Yu-Ting Lin & Ya-Wen Cheng, 2016. "Document overlapping clustering using formal concept analysis," Journal of Advances in Technology and Engineering Research, A/Professor Akbar A. Khatibi, vol. 2(2), pages 28-34.
    6. D. Thorleuchter & D. Van Den Poel, 2013. "Semantic Compared Cross Impact Analysis," Working Papers of Faculty of Economics and Business Administration, Ghent University, Belgium 13/862, Ghent University, Faculty of Economics and Business Administration.
    7. Han Zou & Jing Ge & Ruichao Liu & Lin He, 2023. "Feature Recognition of Regional Architecture Forms Based on Machine Learning: A Case Study of Architecture Heritage in Hubei Province, China," Sustainability, MDPI, vol. 15(4), pages 1-27, February.
    8. Josef Schwaiger & Timo Hammerl & Johannsen Florian & Susanne Leist, 2021. "UR: SMART–A tool for analyzing social media content," Information Systems and e-Business Management, Springer, vol. 19(4), pages 1275-1320, December.
    9. Verwaeren, Jan & Waegeman, Willem & De Baets, Bernard, 2012. "Learning partial ordinal class memberships with kernel-based proportional odds models," Computational Statistics & Data Analysis, Elsevier, vol. 56(4), pages 928-942.
    10. Huazhen Wang & Xin Liu & Bing Lv & Fan Yang & Yanzhu Hong, 2014. "Reliable Multi-Label Learning via Conformal Predictor and Random Forest for Syndrome Differentiation of Chronic Fatigue in Traditional Chinese Medicine," PLOS ONE, Public Library of Science, vol. 9(6), pages 1-14, June.
    11. D. Thorleuchter & D. Van Den Poel & A. Prinzie & -, 2010. "A compared R&D-based and patent-based cross impact analysis for identifying relationships between technologies," Working Papers of Faculty of Economics and Business Administration, Ghent University, Belgium 10/632, Ghent University, Faculty of Economics and Business Administration.
    12. Azzini, Antonia & Cortesi, Nicola & Marrara, Stefania & Topalović, Amir, 2019. "A Multi-Label Machine Learning Approach to Support Pathologist's Histological Analysis," Proceedings of the ENTRENOVA - ENTerprise REsearch InNOVAtion Conference (2019), Rovinj, Croatia, in: Proceedings of the ENTRENOVA - ENTerprise REsearch InNOVAtion Conference, Rovinj, Croatia, 12-14 September 2019, pages 197-208, IRENET - Society for Advancing Innovation and Research in Economy, Zagreb.
    13. Bocheng Li & Yunqiu Zhang & Xusheng Wu, 2022. "DLKN-MLC: A Disease Prediction Model via Multi-Label Learning," IJERPH, MDPI, vol. 19(15), pages 1-15, August.
    14. Francisco J. Ribadas-Pena & Shuyuan Cao & Víctor M. Darriba Bilbao, 2022. "Improving Large-Scale k -Nearest Neighbor Text Categorization with Label Autoencoders," Mathematics, MDPI, vol. 10(16), pages 1-22, August.
    15. Tao Shu & Zhiyi Wang & Huading Jia & Wenjin Zhao & Jixian Zhou & Tao Peng, 2022. "Consumers’ Opinions towards Public Health Effects of Online Games: An Empirical Study Based on Social Media Comments in China," IJERPH, MDPI, vol. 19(19), pages 1-19, October.
    16. Bogaert, Matthias & Lootens, Justine & Van den Poel, Dirk & Ballings, Michel, 2019. "Evaluating multi-label classifiers and recommender systems in the financial service sector," European Journal of Operational Research, Elsevier, vol. 279(2), pages 620-634.
    17. Debaere, Steven & Coussement, Kristof & De Ruyck, Tom, 2018. "Multi-label classification of member participation in online innovation communities," European Journal of Operational Research, Elsevier, vol. 270(2), pages 761-774.

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:wsi:jikmxx:v:15:y:2016:i:01:n:s0219649216500088. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Tai Tone Lim (email available below). General contact details of provider: http://www.worldscinet.com/jikm/jikm.shtml .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.