The largest open Persian speech corpus, with zero-shot voice cloning
@article{ranjbar2025parsvoice,
title={ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis},
author={Ranjbar Kalahroodi, Mohammad Javad and Faili, Heshaam and Shakery, Azadeh},
journal={arXiv preprint arXiv:2510.10774},
year={2025}
}
Existing Persian speech datasets are typically far smaller than their English counterparts, which limits Persian speech technology. ParsVoice addresses this with an automated pipeline that turns raw audiobook recordings into TTS-ready data — a BERT-based sentence-completion detector, a binary-search boundary-optimization step for precise audio-text alignment, and audio-text quality assessment tailored to Persian.
The pipeline processes 2,000 audiobooks into 3,526 hours of clean speech, filtered down to a 1,804-hour high-quality subset spanning 470+ speakers. Fine-tuning XTTS on this data reaches a naturalness MOS of 3.6/5 and a speaker-similarity SMOS of 4.0/5 — comparable to major English corpora, and openly released to accelerate Persian speech research.
Zero-shot clones from the fine-tuned model, next to the unseen reference voice each was cloned from.
Loading samples…