What Web Template Extractor Should I Use? A Benchmarking and Comparison for Five Template Extractors-Reference-Cited by-同舟云学术

What Web Template Extractor Should I Use? A Benchmarking and Comparison for Five Template Extractors

Published:2019-04-12 Issue:2 Volume:13 Page:1-19
ISSN:1559-1131
Container-title:ACM Transactions on the Web
language:en
Short-container-title:ACM Trans. Web

Author:

Alarte Julián¹,Silva Josep¹,Tamarit Salvador²

Affiliation:

1. Universitat Politècnica de València, Spain

2. Universitat Politècnica de Madrid, Spain

Abstract

A Web template is a resource that implements the structure and format of a website, making it ready for plugging content into already formatted and prepared pages. For this reason, templates are one of the main development resources for website engineers, because they increase productivity. Templates are also useful for the final user, because they provide uniformity and a common look and feel for all webpages. However, from the point of view of crawlers and indexers, templates are an important problem, because templates usually contain irrelevant information, such as advertisements, menus, and banners. Processing and storing this information leads to a waste of resources (storage space, bandwidth, etc.). It has been measured that templates represent between 40% and 50% of data on the Web. Therefore, identifying templates is essential for indexing tasks. There exist many techniques and tools for template extraction, but, unfortunately, it is not clear at all which template extractor should a user/system use, because they have never been compared, and because they present different (complementary) features such as precision, recall, and efficiency. In this work, we compare the most advanced template extractors. We implemented and evaluated five of the most advanced template extractors in the literature. To compare all of them, we implemented a workbench, where they have been integrated and evaluated. Thanks to this workbench, we can provide a fair empirical comparison of all methods using the same benchmarks, technology, implementation language, and evaluation criteria.

Funder

Generalitat Valenciana

Spanish Ministerio de Ciencia, Innovacio?n y Universidades/AEI

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications

Link

https://dl.acm.org/doi/pdf/10.1145/3316810

Reference60 articles.

1. TeMex

2. Site-Level Web Template Extraction Based on DOM Analysis

3. Effectiveness of template detection on noise reduction and websites summarization

4. Template detection via data mining and its applications

Cited by 7 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. An Empirical Comparison of Web Content Extraction Algorithms;Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval;2023-07-18

2. A WebExtension framework for experimentation and evaluation of webpage segmentation methods;SoftwareX;2023-07

3. JSAnalyzer: A Web Developer Tool for Simplifying Mobile Web Pages through Non-critical JavaScript Elimination;ACM Transactions on the Web;2022-11-16

4. Rapid Development of a Data Visualization Service in an Emergency Response;IEEE Transactions on Services Computing;2022-05-01

5. HybEx: A Hybrid Tool for Template Extraction;Companion Proceedings of the Web Conference 2022;2022-04-25