z-logo
open-access-imgOpen Access
Improving the Accuracy of the Speech Synthesis Based Phonetic Alignment Using Multiple Acoustic Features
Author(s) -
Sérgio Roberto de Paulo,
Luís Oliveira
Publication year - 2003
Publication title -
lecture notes in computer science
Language(s) - English
Resource type - Book series
SCImago Journal Rank - 0.249
H-Index - 400
eISSN - 1611-3349
pISSN - 0302-9743
ISBN - 3-540-40436-8
DOI - 10.1007/3-540-45011-4_5
Subject(s) - computer science , speech recognition , hidden markov model , sequence (biology) , selection (genetic algorithm) , speech synthesis , signal (programming language) , artificial intelligence , natural language processing , genetics , biology , programming language
The phonetic alignment of the spoken utterances for speech research are commonly performed by HMM-based speech recognizers, in forced alignment mode, but the training of the phonetic segment models requires considerable amounts of annotated data. When no such material is available, a possible solution is to synthesize the same phonetic sequence and align the resulting speech signal with the spoken utterances. However, without a careful choice of acoustic features used in this procedure, it can perform poorly when applied to continuous speech utterances. In this paper we propose a new method to select the best features to use in the alignment procedure for each pair of phonetic segment classes. The results show that this selection considerably reduces the segment boundary location errors.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom