Low-Resourced Text-to-Speech System for Cuoi Cham
Abstract
Text-to-Speech (TTS) has become an important research area due to its applications in education, accessibility tools, digital assistants, and language preservation. Although recent end-to-end neural architectures have significantly improved speech synthesis quality, these advances mainly benefit high-resource languages with abundant speech data and linguistic resources. Developing TTS systems for low-resource languages remains challenging because of limited data and insufficient phonological representations. Cuoi Cham (Tho), a minority language spoken in Nghe An province, represents a typical low-resource scenario and remains underrepresented in modern speech technologies. This paper proposes a Cuoi Cham TTS system based on transfer learning using a pre-trained VITS model. To better capture linguistic characteristics, we introduce a linguistically informed phoneme inventory and a phoneme-aware tokenization strategy that explicitly models syllable structure, including onset, nucleus, coda, and tone. Experimental results demonstrate stable and intelligible speech generation under limited data conditions. Objective and subjective evaluations, including Mel Cepstral Distortion (MCD), spectrogram analysis, and Mean Opinion Score (MOS), indicate that combining transfer learning with linguistically informed phoneme representation is a promising approach for low-resource TTS. This work provides an initial framework for Cuoi Cham speech synthesis and contributes to speech technology development for minority languages in Vietnam.
References
Azizah, K., Adriani, M., & Jatmiko, W. (2020). Hierarchical transfer learning for multilingual, multi-speaker, and style transfer DNN-based TTS on low-resource languages. IEEE Access, 8, 179798–179812. https://doi.org/10.1109/ACCESS.2020.3027619.
Byambadorj, Z., Nishimura, R., Ayush, A., Ohta, K., & Kitaoka, N. (2021). Text-to-speech system for low-resource language using cross-lingual transfer learning and data augmentation. EURASIP Journal on Audio, Speech, and Music Processing, 2021(1), 42. https://doi.org/10.1186/s13636-021-00225-4.
Casanova, E., Weber, J., Shulby, C. D., Junior, A. C., Gölge, E., & Ponti, M. A. (2022). YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone. In Proceedings of the 39th International Conference on Machine Learning (pp. 2709–2720). PMLR. https://proceedings.mlr.press/v162/casanova22a.html.
Hwang, M.-J., Yamamoto, R., Song, E., & Kim, J.-M. (2020). TTS-by-TTS: TTS-driven data augmentation for fast and high-quality speech synthesis. arXiv:2010.13421. https://doi.org/10.48550/arXiv.2010.13421.
Joshi, R., & Garera, N. (2023). Rapid speaker adaptation in low resource text to speech systems using synthetic data and transfer learning. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation (pp. 267–273).
Kim, J., Kong, J., & Son, J. (2021). Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In Proceedings of the 38th International Conference on Machine Learning (Vol. 139, pp. 5530–5540).
Kubichek, R. (1993). Mel-cepstral distance measure for objective speech quality assessment. In Proceedings of the IEEE Pacific Rim Conference on Communications, Computers and Signal Processing (Vol. 1, pp. 125–128). https://doi.org/10.1109/PACRIM.1993.407206.
Liu, R., Wen, X., Lu, C., & Chen, X. (2020). Tone learning in low-resource bilingual TTS. In Interspeech 2020 (pp. 2952–2956). ISCA. https://doi.org/10.21437/Interspeech.2020-2180.
Michailovsky, B., Mazaudon, M., Michaud, A., Guillaume, S., François, A., & Adamou, E. (2014). Documenting and researching endangered languages: The Pangloss Collection. Language Documentation & Conservation, 8, 119–135.
Shen, J., et al. (2018). Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions. arXiv:1712.05884. https://doi.org/10.48550/arXiv.1712.05884.
Tan, X., Qin, T., Soong, F., & Liu, T.-Y. (2021). A survey on neural speech synthesis. arXiv:2106.15561. https://doi.org/10.48550/arXiv.2106.15561.
van den Oord, A., et al. (2016). WaveNet: A generative model for raw audio. arXiv:1609.03499. https://doi.org/10.48550/arXiv.1609.03499.
Wang, Y., et al. (2017). Tacotron: Towards end-to-end speech synthesis. arXiv:1703.10135. https://doi.org/10.48550/arXiv.1703.10135.
Copyright (c) 2026 Dong Pham Van, Long Vu Duy, Hai Nguyen Duy

This work is licensed under a Creative Commons Attribution 4.0 International License.
ISSN 

