PersianVox: A Prosody-Aware Dual-ASR Approach for Speech Corpus Generation from In-the-Wild Data
Published in arXiv, 2026
Abstract: The advancement of zero-shot text-to-speech (TTS) synthesis is currently hindered for low-resource languages by the scarcity of large-scale, high-fidelity speech datasets. Traditional alignment-based methods require rare verbatim transcripts, while standard in-the-wild pipelines often rely on single-model automatic speech recognition (ASR) and silence-based segmentation, leading to transcription errors and truncated prosody. To address these challenges for the Persian language, this paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabelled web data. Our approach integrates a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling. Furthermore, we employ a dual-ASR agreement mechanism, leveraging two distinct model architectures to filter unreliable transcriptions without ground truth. This pipeline yields a 1600-hour multi-speaker dataset, the largest open-source speech resource available for Persian to date. Additionally, we provide the first comparative benchmark of Speech Quality Assessment (SQA) methods for Persian, releasing a human-annotated subset to facilitate future research.
Recommended citation: Zouashkiani, S., Khalesi, S., Soleimani Roudi, S., Amini, S., & Ghaemmaghami, S. (2025). "PersianVox: A Prosody-Aware Dual-ASR Approach for Speech Corpus Generation from In-the-Wild Data." arXiv.
Download Paper
