Vietnamese Non-Verbal Text-to-Speech for Conversational Synthesis
Important dates
September 23: Training Data release
September 30: Public Test
October 9: Private Test
October 15: Result announcement
October 25: Paper submission
November 5: Acceptance notification
November 12: Camera-ready (=> Proceedings)
November 15: Workshop date
General Description
ViNV-TTS is a shared task on synthesizing natural Vietnamese conversational speech that contains non-verbal vocalizations (NVs) - communicative sounds that are not words, such as laughter, breathing, sniffing, and throat-clearing. Given Vietnamese text with inline NV tags at explicit positions, systems must produce speech in which each tagged NV is present, of the correct type, placed at the intended position, and acoustically natural and well-integrated with the surrounding speech, while preserving intelligibility and the identity of the target speaker.
The task has two tag-conditioned tracks that differ only in speaker setting:
• Track A (Core): synthesizes speech for a target speaker seen during training. Zero-shot voice cloning is not required.
• Track B (Advanced): synthesizes speech for an unseen target speaker given a short reference audio clip, requiring zero-shot voice cloning.
Systems are evaluated through both human and automatic evaluation. Human evaluation includes speech naturalness, audio quality, NV naturalness and NV placement accuracy. Automatic evaluation includes WER, predicted MOS and speaker similarity.
Dataset
The task releases the first public Vietnamese conversational speech corpus annotated with non-verbal tags, providing resources for training and system development. It is built from Vietnamese conversational sources such as podcasts and talk shows via an automated construction pipeline.
The corpus is multi-speaker, segmented into short conversational utterances, and annotated with four non-verbal tags - [laughter], [breathing], [sniff], and [throatclearing] - placed inline at their point of occurrence. Each item provides the audio (single-channel, 24 kHz), a transcript containing inline non-verbal tags and a speaker ID. Exact corpus statistics (total duration, number of utterances, and per-tag counts) will be finalized and reported at release.
The data is split into public Train/Dev and Public test portions, plus a held-out Private test with two versions:
• Track A: new text with gold non-verbal tags on speakers seen during training
• Track B: new text with gold non-verbal tags and short reference audio clips from held-out, unseen speakers, requiring zero-shot voice cloning.
Evaluation Metrics
Three main aspects will be used to evaluate the submitted models: Non-Verbal Vocalization (NV) Realization, Speech/Audio Quality, and Speaker Similarity, using both automatic and human evaluation methods.
• NV Placement Accuracy (NVPA): Measures whether the required non-verbal vocalizations are correctly realized at their intended positions in the synthesized speech. An NV event is considered correctly realized only when the required NV type is present at the intended position. Range: 0 to 1, where higher is better.
• Speech Naturalness (SN): Measures the overall naturalness, fluency, and human-likeness of the synthesized speech, including the speech surrounding the non-verbal vocalizations. Scale: 1 to 5, where higher is better.
• Quality (Q): Measures the overall perceptual quality of the synthesized audio, including audible artifacts, distortions, noise, and other audio degradations. Scale: 1 to 5, where higher is better.
• Word Error Rate (WER): Evaluates the intelligibility and content accuracy of the synthesized speech by comparing the recognized speech with the input transcript using a pretrained ASR model (Zipformer). Non-verbal tags are excluded from WER computation. Range: 0 to 1, where lower is better.
• Predicted MOS (pMOS): Measures the predicted perceptual quality and naturalness of the synthesized speech using an automatic MOS prediction model (DNSMOS). Scale: 1 to 5, where higher is better.
• Speaker Similarity (SS): Measures the similarity between the target/reference speaker and the synthesized speech using a pretrained speaker recognition model (ECAPA-TDNN). Range: 0 to 1, where higher is better.
Final Scoring Formula:
• Track A: A = 0.3*NVPA + 0.15*SN + 0.15*Q + 0.15*(1 - WER) + 0.15*pMOS + 0.1*SS;
• Track B: B = 0.3*NVPA + 0.15*SN + 0.15*Q + 0.1*(1 - WER) + 0.1*pMOS + 0.2*SS
*Note: All components of the scoring formula are normalized to [0, 1].
Submission Format
After training their systems, participating teams are required to deploy their models as an API endpoint. During the evaluation phase, the organizers will send inference requests to the submitted API and collect the resulting synthesized speech. Each team will be provided with an allocated evaluation period, which will be announced by the organizers in advance. The API must remain accessible and operational throughout the assigned evaluation period and return the synthesized audio for each valid input request.
Contest rules
- External Data and Synthetic Data: Teams are allowed to use external datasets and synthetic data generated by TTS models, provided that the data and models used are open-source. Teams must report the sources of all external data and the models used to generate synthetic data. Registration deadline: End of October 3, 2026.
- Pre-trained Models and External Services: Teams are allowed to use publicly available pre-trained TTS models. However, the use of external APIs or external TTS services to generate the submitted results is strictly prohibited. Any pre-trained models used by a team must be reported to the Organizing Committee for approval. The approved models will be publicly announced to all participating teams. Registration deadline: End of October 3, 2026.
- Reproducibility Requirement: Teams whose results are ranked in the Top 5 must submit their source code to the Organizing Committee for result reproduction and verification, in order to ensure fairness and transparency of the competition.
- Technical Report Requirement: Award-winning teams are required to submit a technical report describing their approach, including the selected models and datasets, training/fine-tuning/inference setup, and experimental results (if any). Failure to submit the required technical report will result in the competition result being deemed invalid and not officially recognized by the Organizing Committee.
Contact
Zalo Group: ViNV-TTS for Conversational Synthesis
Registration
https://docs.google.com/forms/d/e/1FAIpQLSdT2KQCWZAsb3BYgqfD2uBDXg_SFPt5LN3cUmeHrZ46fZqb6Q/viewform
Organizers
- Nguyễn Hoàng Kỳ - Vbee AITalk JSC - kynh@vbeeai.com
- Trần Sơn Tùng - Vbee AITalk JSC - tungts@vbeeai.com
- Nguyễn Thị Thu Trang
Vbee AITalk JSC - trangntt@vbeeai.com
Hanoi University of Science and Technology (HUST) - trangntt@soict.hust.edu.vn
References
- Yao, Z., Guo, L., Yang, X., Kang, W., Kuang, F., Yang, Y., ... & Povey, D. (2024, May). Zipformer: A faster and better encoder for automatic speech recognition. In International Conference on Learning Representations (Vol. 2024, pp. 44440-44455).
- Ye, R., Zhou, Y., Yu, R., Lin, Z., Li, K., Li, X., ... & Wu, Z. (2025). A scalable pipeline for enabling non-verbal speech generation and understanding. arXiv preprint arXiv:2508.05385.



