z-logo
open-access-imgOpen Access
Scalable browsing for large collections
Author(s) -
Gordon W. Paynter,
Ian H. Witten,
Sally Jo Cunningham,
George Buchanan
Publication year - 2000
Publication title -
research commons (university of waikato)
Language(s) - English
Resource type - Conference proceedings
ISBN - 1-58113-231-X
DOI - 10.1145/336597.336666
Subject(s) - computer science , phrase , scalability , interface (matter) , simple (philosophy) , hierarchy , information retrieval , natural language processing , artificial intelligence , world wide web , database , epistemology , bubble , philosophy , maximum bubble pressure method , parallel computing , market economy , economics
Phrase browsing techniques use phrases extracted automatically from a large information collection as a basis for browsing and accessing it. This paper describes a case study that uses an automatically constructed phrase hierarchy to facilitate browsing of an ordinary large Web site. Phrases are extracted from the full text using a novel combination of rudimentary syntactic processing and sequential grammar induction techniques. The interface is simple, robust and easy to use. To convey a feeling for the quality of the phrases that are generated automatically, a thesaurus used by the organization responsible for the Web site is studied and its degree of overlap with the phrases in the hierarchy is analyzed. Our ultimate goal is to amalgamate hierarchical phrase browsing and hierarchical thesaurus browsing: the latter provides an authoritative domain vocabulary and the former augments coverage in areas the thesaurus does not reach

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom