Practical NLP-Based Text Indexing
Author(s) -
Jesús Vilares,
Fco. Mario Barcala,
Miguel Á. Alonso,
Jorge Graña
Publication year - 2002
Publication title -
lecture notes in computer science
Language(s) - English
Resource type - Book series
SCImago Journal Rank - 0.249
H-Index - 400
eISSN - 1611-3349
pISSN - 0302-9743
ISBN - 3-540-00131-X
DOI - 10.1007/3-540-36131-6_65
Subject(s) - computer science , search engine indexing , parsing , natural language processing , artificial intelligence , dependency (uml) , set (abstract data type) , dependency grammar , automatic indexing , text processing , information retrieval , programming language
We consider a set of natural language processing techniques based on finite-state technology that can be used to analyze huge amounts of texts. These techniques include an advanced tokenizer, a part-of-speech tagger that can manage ambiguous streams of words, a system for conflating words by means of derivational mechanisms, and a shallow parser to extract syntactic-dependency pairs. We propose to use these techniques in order to improve the performance of standard indexing engines.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom