Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (Contributo in atti di convegno)

Type
Label
  • Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (Contributo in atti di convegno) (literal)
Anno
  • 2011-01-01T00:00:00+01:00 (literal)
Alternative label
  • Zito C., Tesser F., Nicolao M., Cosi P. (2011)
    Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech
    in AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, "Contesto comunicativo e variabilità nella produzione e percezione della lingua", Università del Salento - Lecce, 26-28 gennaio 2011
    (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#autori
  • Zito C., Tesser F., Nicolao M., Cosi P. (literal)
Pagina inizio
  • 392 (literal)
Pagina fine
  • 403 (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#altreInformazioni
  • Zito C., Tesser F., Nicolao M., Cosi P. \"Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech\" in Abstract Book & CD-Rom Proceedings of AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\" 26-28 gennaio 2011, Università del Salento - Lecce Abstract Book: 81 - (CD: 392-403). (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#url
  • http://www.cs.bham.ac.uk/~cxz004/download/paper/2011/zito-tesser-nicolao-cosi.pdf (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#titoloVolume
  • AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\" (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#volumeInCollana
  • 7 (1) (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#affiliazioni
  • Claudio Zito - Dipartimento di Informatica, Università di Pisa, Italia Fabio Tesser, Piero Cosi -Istituto di Scienze e Tecnologie della Cognizione, Consiglio Nazionale di Ricerca, Italia Mauro Nicolao - Speech and Hearing Research Group, University of Sheffield, United Kingdom (literal)
Titolo
  • Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#isbn
  • 978-88-7870-619-4 (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#autoriVolume
  • \"Statistical Context-Dependent Units Boundary Correction for Corpus-Based Unit-Selection Text-to-Speech\" in Abstract Book & CD-Rom Proceedings of AISV 2011, 7th Conference of Associazione Italiana di Scienze della Voce, \"Contesto comunicativo e variabilità nella produzione e percezione della lingua\", 26-28 gennaio 2011, Università del Salento - Lecce - Abs.: 81 - (CD: 392-403). (literal)
Http://www.cnr.it/ontology/cnr/pubblicazioni.owl#curatoriVolume
  • B. Gili Fivela, A. Stella, L. Garrapa, M. Grimaldi (literal)
Abstract
  • In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques for speaker adaptation, which attempt to improve the accuracy of the segmentation using acoustic models that are more robust in the face of the speaker's characteristics, we aim to use only context dependent characteristics extrapolated with linguistic analysis techniques. In simple terms, we use the intuitive idea that context dependent information is tightly correlated with the related acoustic waveform. We propose a statistical model, which predicts correcting values to reduce the systematic error produced by a state-of-the-art Hidden Markov Model (HMM) based speech segmentation. In other words, we can predict how HMM-based Automatic Speech Recognition (ASR) systems interpret the waveform signal determining the systematic error in different contextual scenarios. Our approach consists of two phases: (1) identifying context-dependent phonetic unit classes (for instance, the class which identifies vowels as being the nucleus of monosyllabic words); and (2) building a regression model that associates the mean error value made by the ASR during the segmentation of a single speaker corpus to each class. The success of the approach is evaluated by comparing the corrected boundaries of units and the state-of-the-art HHM segmentation against a reference alignment, which is supposed to be the optimal solution. The results of this study show that the context-dependent correction of units' boundaries has a positive influence on the forced alignment, especially when the misinterpretation of the phone is driven by acoustic properties linked to the speaker's phonetic characteristics. In conclusion, our work supplies a first analysis of a model sensitive to speaker-dependent characteristics, robust to defective and noisy information, and a very simple implementation which could be utilized as an alternative to either more expensive speaker-adaptation systems or of numerous manual correction sessions. (literal)
Editore
Prodotto di
Autore CNR
Insieme di parole chiave

Incoming links:


Autore CNR di
Prodotto
Editore di
Insieme di parole chiave di
data.CNR.it