A General Approach for Partitioning Web Page Content Based on Geometric and Style Information
Author(s) -
Hai-Feng Guo,
Jalal Mahmud,
Yevgen Borodin,
Amanda Stent,
I. V. Ramakrishnan
Publication year - 2007
Publication title -
ninth international conference on document analysis and recognition (icdar 2007)
Language(s) - English
Resource type - Book series
ISBN - 0-7695-2822-8
DOI - 10.1109/icdar.2007.10
In this paper, we describe a general-purpose approach for partitioning Web page content. The novelty of our ap- proach lies in the use of detailed layout information from a Web page renderer to determine spatial locality and identify visual separators, and the use of relaxed matching over pre- sentation style information to determine presentation style similarity. We present several examples to illustrate the gen- erality of our approach.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom