Author
Listed:
- Anil Kumar Sinha
(Department of Computer Science, V.K.S. University Ara, India)
- Khushboo Mishra
(P.G. Department of Physics, V.K.S. University, Ara , India)
- Md Alimul Haque
(Department of Computer Science, V.K.S. University Ara, India)
- B. K. Mishra
(Department of Computer Science, V.K.S. University Ara, India)
Abstract
With the exponential growth of web-based content, efficient retrieval of contextually relevant textual information starting from seed URLs has become a critical challenge in web content mining and information retrieval. Traditional crawling and search methods—such as breadth-first search (BFS), depth-first search (DFS), best-first (focused crawling), topic-sensitive PageRank, and context-graph models—typically suffer from limitations such as parameter tuning overhead, lack of contextual understanding, requirement of large training datasets, high computational cost, and the need for specialised infrastructure. This research presents a comprehensive comparative study of multiple search and crawling models applied to textual retrieval from seed URLs, with a particular focus on their performance in diverse web‐structures (static vs dynamic) and content types. Employing a unified experimental framework implemented in Python with MySQL backend, we evaluate each algorithm using standard performance metrics (precision, recall, F1-score) alongside newer metrics such as coverage, relevance score, search time, memory usage, throughput and harvest rate. Machine-learning enabled variants (for example semantic-BFS and semantic-DFS using transformer-based embeddings) are also incorporated to assess their value over purely structural methods. Our results demonstrate that while semantic-enhanced BFS (Semantic-BFS) yields higher coverage, better relevance and faster response time in many scenarios, it shows limitations in classical metrics like precision/recall/F1 when ground-truth labels are inadequate for semantic relevance. The study provides insights into algorithmic trade-offs, suitability for different web architectures, and proposes hybrid strategies for next-generation crawlers and retrieval systems. The findings contribute toward the design of more adaptive, semantic-aware, and scalable web content mining frameworks.
Suggested Citation
Anil Kumar Sinha & Khushboo Mishra & Md Alimul Haque & B. K. Mishra, 2026.
"Performance Analysis of Transformer-Enabled Semantic Crawlers for Scalable Text Retrieval,"
Diginomics, Centro de Estudios de Economía Digital, vol. 5, pages 309-309, January.
Handle:
RePEc:cwg:digino:v:5:y:2026:id:309
DOI: 10.56294/digi2026309
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:cwg:digino:v:5:y:2026:id:309. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Prof. Carlos Alberto Gómez Cano (email available below). General contact details of provider: https://diginomics.ar/index.php/digi .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.