z-logo
open-access-imgOpen Access
Measuring Contribution of HTML Features in Web Document Clustering
Author(s) -
Esteban Meneses,
Oldemar Rodríguez-Rojas
Publication year - 2008
Publication title -
clei electronic journal
Language(s) - English
Resource type - Journals
ISSN - 0717-5000
DOI - 10.19153/cleiej.11.2.7
Subject(s) - computer science , document clustering , suffix , cluster analysis , information retrieval , partition (number theory) , feature (linguistics) , representation (politics) , html element , term (time) , data mining , artificial intelligence , web page , world wide web , mathematics , linguistics , combinatorics , law , political science , quantum mechanics , philosophy , physics , politics
Documents in HTML format have many features to analyze, from the terms in special sections to the phrases that appear in the whole document. However, it is important to decide which feature contributes the most to separate documents according to classes. Given this information, it is possible not to include certain feature in the representation for the document, given that it is expensive to compute and doesn’t contribute enough in the clustering process. By using a novel representation model and the standard k-means algorithm, we discovered that terms in the body of document contributes the most, followed by terms in other sections. Suffix tree provides poor contribution in that scenario, while term order graphs influence a little the partition. We used 4 known datasets to support the conclusions.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom