z-logo
open-access-imgOpen Access
Exploiting content redundancy for web information extraction
Author(s) -
Pankaj Gulhane,
Rajeev Rastogi,
Srinivasan H. Sengamedu,
Ashwin Tengli
Publication year - 2010
Publication title -
proceedings of the vldb endowment
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.946
H-Index - 134
ISSN - 2150-8097
DOI - 10.14778/1920841.1920915
Subject(s) - computer science , redundancy (engineering) , data mining , exploit , information retrieval , similarity (geometry) , a priori and a posteriori , matching (statistics) , web page , metric (unit) , set (abstract data type) , artificial intelligence , mathematics , world wide web , image (mathematics) , philosophy , statistics , operations management , computer security , epistemology , economics , programming language , operating system
We propose a novel extraction approach that exploits content redundancy on the web to extract structured data from template-based web sites. We start by populating a seed database with records extracted from a few initial sites. We then identify values within the pages of each new site that match attribute values contained in the seed set of records. To match attribute values with diverse representations across sites, we define a new similarity metric that leverages the templatized structure of attribute content. Specifically, our metric discovers the matching pattern between attribute values from two sites, and uses this to ignore extraneous portions of attribute values when computing similarity scores. Further, to filter out noisy attribute value matches, we exploit the fact that attribute values occur at fixed positions within template-based sites. We develop an efficient Apriori-style algorithm to systematically enumerate attribute position configurations with sufficient matching values across pages. Finally, we conduct an extensive experimental study with real-life web data to demonstrate the effectiveness of our extraction approach.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom