A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models

kirill, Nikita, Vasiliy kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Oleg Rogov, Grach Mkrtchian

A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models: 53 upvotes on Hugging Face Daily Papers, #2 of 11 papers on 2025-07-21. Day-by-day upvote history.

Russian speech synthesis presents distinctive challenges, including vowel reduction, consonant devoicing, variable stress patterns, homograph ambiguity, and unnatural intonation. This paper introduces Balalaika, a novel dataset comprising more than 2,000 hours of studio-quality Russian speech with comprehensive textual annotations, including punctuation and stress markings. Experimental results show that models trained on Balalaika significantly outperform those trained on existing datasets in both speech synthesis and enhancement tasks. We detail the dataset construction pipeline, annotation methodology, and results of comparative evaluations.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.