http://www.cnr.it/ontology/cnr/individuo/prodotto/ID239602
Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (Contributo in atti di convegno)
- Type
- Label
- Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (Contributo in atti di convegno) (literal)
- Anno
- 2011-01-01T00:00:00+01:00 (literal)
- Alternative label
Zito C., Tesser F., Nicolao M., Cosi P. (2011)
Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech
in AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, "Contesto comunicativo e variabilità nella produzione e percezione della lingua", Università del Salento - Lecce, 26-28 gennaio 2011
(literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#autori
- Zito C., Tesser F., Nicolao M., Cosi P. (literal)
- Pagina inizio
- Pagina fine
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#altreInformazioni
- Zito C., Tesser F., Nicolao M., Cosi P.
\"Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech\"
in Abstract Book & CD-Rom Proceedings of
AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\"
26-28 gennaio 2011, Università del Salento - Lecce
Abstract Book: 81 - (CD: 392-403). (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#url
- http://www.cs.bham.ac.uk/~cxz004/download/paper/2011/zito-tesser-nicolao-cosi.pdf (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#titoloVolume
- AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\" (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#volumeInCollana
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#affiliazioni
- Claudio Zito - Dipartimento di Informatica, Università di Pisa, Italia
Fabio Tesser, Piero Cosi -Istituto di Scienze e Tecnologie della Cognizione, Consiglio Nazionale di Ricerca, Italia
Mauro Nicolao - Speech and Hearing Research Group, University of Sheffield, United Kingdom (literal)
- Titolo
- Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#isbn
- 978-88-7870-619-4 (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#autoriVolume
- \"Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech\"
in Abstract Book & CD-Rom Proceedings of AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\", 26-28 gennaio 2011, Università del Salento - Lecce - Abs.: 81 - (CD: 392-403). (literal)
- Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#curatoriVolume
- B. Gili Fivela, A. Stella, L. Garrapa, M. Grimaldi (literal)
- Abstract
- In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques for speaker adaptation, which attempt to improve the accuracy of the segmentation using acoustic models that are more robust in the face of the speaker's characteristics, we aim to use only context dependent characteristics extrapolated with linguistic analysis techniques. In simple terms, we use the intuitive idea that context dependent information is tightly correlated with the related acoustic waveform. We propose a statistical model, which predicts correcting values to reduce the systematic error produced by a state-of-the-art Hidden Markov Model (HMM) based speech segmentation. In other words, we can predict how HMM-based Automatic Speech Recognition (ASR) systems interpret the waveform signal determining the systematic error in different contextual scenarios. Our approach consists of two phases: (1) identifying context-dependent phonetic unit classes (for instance, the class which identifies vowels as being the nucleus of monosyllabic words); and (2) building a regression model that associates the mean error value made by the ASR during the segmentation of a single speaker corpus to each class. The success of the approach is evaluated by comparing the corrected boundaries of units and the state-of-the-art HHM segmentation against a reference alignment, which is supposed to be the optimal solution. The results of this study show that the context-dependent correction of units' boundaries has a positive influence on the forced alignment, especially when the misinterpretation of the phone is driven by acoustic properties linked to the speaker's phonetic characteristics. In conclusion, our work supplies a first analysis of a model sensitive to speaker-dependent characteristics, robust to defective and noisy information, and a very simple implementation which could be utilized as an alternative to either more expensive speaker-adaptation systems or of numerous manual correction sessions. (literal)
- Editore
- Prodotto di
- Autore CNR
- Insieme di parole chiave
Incoming links:
- Autore CNR di
- Prodotto
- Editore di
- Insieme di parole chiave di