IHWC: intelligent hidden web crawler for harvesting data in urban domains-Reference-Cited by-同舟云学术

IHWC: intelligent hidden web crawler for harvesting data in urban domains

Published:2021-07-24 Issue: Volume: Page:
ISSN:2199-4536
Container-title:Complex & Intelligent Systems
language:en
Short-container-title:Complex Intell. Syst.

Author:

Kaur Sawroop,Singh Aman^ORCID,Geetha G.,Cheng Xiaochun

Abstract

AbstractDue to the massive size of the hidden web, searching, retrieving and mining rich and high-quality data can be a daunting task. Moreover, with the presence of forms, data cannot be accessed easily. Forms are dynamic, heterogeneous and spread over trillions of web pages. Significant efforts have addressed the problem of tapping into the hidden web to integrate and mine rich data. Effective techniques, as well as application in special cases, are required to be explored to achieve an effective harvest rate. One such special area is atmospheric science, where hidden web crawling is least implemented, and crawler is required to crawl through the huge web to narrow down the search to specific data. In this study, an intelligent hidden web crawler for harvesting data in urban domains (IHWC) is implemented to address the relative problems such as classification of domains, prevention of exhaustive searching, and prioritizing the URLs. The crawler also performs well in curating pollution-related data. The crawler targets the relevant web pages and discards the irrelevant by implementing rejection rules. To achieve more accurate results for a focused crawl, ICHW crawls the websites on priority for a given topic. The crawler has fulfilled the dual objective of developing an effective hidden web crawler that can focus on diverse domains and to check its integration in searching pollution data in smart cities. One of the objectives of smart cities is to reduce pollution. Resultant crawled data can be used for finding the reason for pollution. The crawler can help the user to search the level of pollution in a specific area. The harvest rate of the crawler is compared with pioneer existing work. With an increase in the size of a dataset, the presented crawler can add significant value to emission accuracy. Our results are demonstrating the accuracy and harvest rate of the proposed framework, and it efficiently collect hidden web interfaces from large-scale sites and achieve higher rates than other crawlers.

Publisher

Springer Science and Business Media LLC

Subject

General Earth and Planetary Sciences,General Environmental Science

Link

https://link.springer.com/content/pdf/10.1007/s40747-021-00471-1.pdf

Reference45 articles.

1. Kobayashi M, Takeda K (2000) Information retrieval on the Web. ACM Comput Surv 32(2):144–173

2. Wu M, Lee C (2020) A study on natural language processing classified news. pp. 244–247

3. Chakrabarti S (2003) Crawling the web. In: Mining the web, pp 17–43

4. Kaur S, Geetha G (2007) Advances in web crawlers, vol. 10, pp. 1–22

5. Li Y, Wang Y, Du J (2013) E-FFC: an enhanced form-focused crawler for domain-specific deep web databases. J Intell Inform Syst 40(1):159–184

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Bot crawler to retrieve data from Facebook based on the selection of posts and the extraction of user profiles;Inge CuC;2022-09-20

2. Correction: IHWC: intelligent hidden web crawler for harvesting data in urban domains;Complex & Intelligent Systems;2022-08-10

3. Effect Analysis of Carbon Information on Enterprise Value Based on Big Data;Mathematical Problems in Engineering;2022-07-11

4. Internet Driven Dynamic Question Analysis and Response for Engineering;2022 45th Jubilee International Convention on Information, Communication and Electronic Technology (MIPRO);2022-05-23