Design and Implementation of Domain based Semantic Hidden Web Crawler

Manvi; Komal Kumar Bhatia; Ashutosh Dixit

arXiv:1509.06847·cs.IR·September 24, 2015·5 cites

Design and Implementation of Domain based Semantic Hidden Web Crawler

Manvi, Komal Kumar Bhatia, Ashutosh Dixit

PDF

Open Access

TL;DR

This paper presents a domain-specific semantic web crawler that effectively extracts hidden web data behind search forms by filling them using semantic mapping to domain databases, improving data retrieval accuracy.

Contribution

It introduces a novel technique that uses semantic mapping to automate form filling for hidden web data extraction, enhancing traditional crawling methods.

Findings

01

Improved accuracy in hidden web data extraction.

02

Effective domain-specific form filling using semantic mapping.

03

Enhanced retrieval of data behind search forms.

Abstract

Web is a wide term which mainly consists of surface web and hidden web. One can easily access the surface web using traditional web crawlers, but they are not able to crawl the hidden portion of the web. These traditional crawlers retrieve contents from web pages, which are linked by hyperlinks ignoring the information hidden behind form pages, which cannot be extracted using simple hyperlink structure. Thus, they ignore large amount of data hidden behind search forms. This paper emphasizes on the extraction of hidden data behind html search forms. The proposed technique makes use of semantic mapping to fill the html search form using domain specific database. Using semantics to fill various fields of a form leads to more accurate and qualitative data extraction.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsWeb Data Mining and Analysis · Caching and Content Delivery · Web visibility and informetrics