Probe, count, and classify-Reference-Cited by-同舟云学术

Probe, count, and classify

Published:2001-06 Issue:2 Volume:30 Page:67-78
ISSN:0163-5808
Container-title:ACM SIGMOD Record
language:en
Short-container-title:SIGMOD Rec.

Author:

Ipeirotis Panagiotis G.¹,Gravano Luis¹,Sahami Mehran²

Affiliation:

1. Computer Science Dept., Columbia University

2. E.piphany, Inc.

Abstract

The contents of many valuable web-accessible databases are only accessible through search interfaces and are hence invisible to traditional web “crawlers.” Recent studies have estimated the size of this “hidden web” to be 500 billion pages, while the size of the “crawlable” web is only an estimated two billion pages. Recently, commercial web sites have started to manually organize web-accessible databases into Yahoo!-like hierarchical classification schemes. In this paper, we introduce a method for automating this classification process by using a small number of query probes. To classify a database, our algorithm does not retrieve or inspect any documents or pages from the database, but rather just exploits the number of matches that each query probe generates at the database in question. We have conducted an extensive experimental evaluation of our technique over collections of real documents, including over one hundred web-accessible databases. Our experiments show that our system has low overhead and achieves high classification accuracy across a variety of databases.

Publisher

Association for Computing Machinery (ACM)

Subject

Information Systems,Software

Link

https://dl.acm.org/doi/pdf/10.1145/376284.375671

Reference32 articles.

1. Automated learning of decision rules for text categorization

2. Automatic discovery of language models for text databases

3. The TESTING OF INDEX LANGUAGE DEVICES

4. Server selection on the World Wide Web

Cited by 21 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Hierarchical confusion matrix for classification performance evaluation;Journal of the Royal Statistical Society Series C: Applied Statistics;2023-07-03

2. Empathic Responses of Behavioral-Synchronization in Human-Agent Interaction;Computers, Materials & Continua;2022

3. SmartCrawler: A Three-Stage Ranking Based Web Crawler for Harvesting Hidden Web Sources;Computers, Materials & Continua;2021

4. Modeling and predicting the user next input by Bayesian reasoning;Soft Computing;2015-10-01

5. Focused crawling for the hidden web;World Wide Web;2015-05-21