A tree-based learning approach for document structure analysis and its application to web search-Reference-Cited by-同舟云学术

A tree-based learning approach for document structure analysis and its application to web search

Published:2014-02-24 Issue:4 Volume:21 Page:569-605
ISSN:1351-3249
Container-title:Natural Language Engineering
language:en
Short-container-title:Nat. Lang. Eng.

Author:

PEMBE F. CANAN,GÜNGÖR TUNGA

Abstract

AbstractIn this paper, we study the problem of structural analysis of Web documents aiming at extracting the sectional hierarchy of a document. In general, a document can be represented as a hierarchy of sections and subsections with corresponding headings and subheadings. We developed two machine learning models: heading extraction model and hierarchy extraction model. Heading extraction was formulated as a classification problem whereas a tree-based learning approach was employed in hierarchy extraction. For this purpose, we developed an incremental learning algorithm based on support vector machines and perceptrons. The models were evaluated in detail with respect to the performance of the heading and hierarchy extraction tasks. For comparison, a baseline rule-based approach was used that relies on heuristics and HTML document object model tree processing. The machine learning approach, which is a fully automatic approach, outperformed the rule-based approach. We also analyzed the effect of document structuring on automatic summarization in the context of Web search. The results of the task-based evaluation on TREC queries showed that structured summaries are superior to unstructured summaries both in terms of accuracy and user ratings, and enable the users to determine the relevancy of search results more accurately than search engine snippets.

Publisher

Cambridge University Press (CUP)

Subject

Artificial Intelligence,Linguistics and Language,Language and Linguistics,Software

Reference54 articles.

1. Web page title extraction and its application

2. Accordion summarization for end-game browsing on PDAs and cellular phones

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Extracting Variable-Depth Logical Document Hierarchy from Long Documents: Method, Evaluation, and Application;Journal of Computer Science and Technology;2022-05-31

2. Differentiating the learning styles of college students in different disciplines in a college English blended learning setting;PLOS ONE;2021-05-20

3. CESS-A System to Categorize Bangla Web Text Documents;ACM Transactions on Asian and Low-Resource Language Information Processing;2020-09-30

4. Automatic categorization of web text documents using fuzzy inference rule;Sādhanā;2020-06-27