
Spoken Document Retrieval Based on Confusion Network with Syllable Fragments
Author(s) -
Lei Zhang,
Yoshihiko Gotoh,
Muhammad Usman Ghani Khan
Publication year - 2012
Publication title -
international journal of advanced robotic systems
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.394
H-Index - 46
eISSN - 1729-8814
pISSN - 1729-8806
DOI - 10.5772/52454
Subject(s) - computer science , sentence , speech recognition , syllable , confusion , word error rate , task (project management) , word (group theory) , document retrieval , artificial intelligence , natural language processing , mathematics , psychology , psychoanalysis , geometry , management , economics
This paper addresses the problem of spoken document retrieval under noisy conditions by incorporating sound selection of a basic unit and an output form of a speech recognition system. Syllable fragment is combined with a confusion network in a spoken document retrieval task. After selecting an appropriate syllable fragment, a lattice is converted into a confusion network that is able to minimize the word error rate instead of maximizing the whole sentence recognition rate. A vector space model is adopted in the retrieval task where tf-idf weights are derived from the posterior probability. The confusion network with syllable fragments is able to improve the mean of average precision (MAP) score by 0.342 and 0.066 over one-best scheme and the lattice