ParsVoice

The largest open Persian speech corpus, with zero-shot voice cloning

1,804h · 470+ speakers Under review, EMNLP 2026
Paper
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
University of Tehran

About ParsVoice

Existing Persian speech datasets are typically far smaller than their English counterparts, which limits Persian speech technology. ParsVoice addresses this with an automated pipeline that turns raw audiobook recordings into TTS-ready data — a BERT-based sentence-completion detector, a binary-search boundary-optimization step for precise audio-text alignment, and audio-text quality assessment tailored to Persian.

The pipeline processes 2,000 audiobooks into 3,526 hours of clean speech, filtered down to a 1,804-hour high-quality subset spanning 470+ speakers. Fine-tuning XTTS on this data reaches a naturalness MOS of 3.6/5 and a speaker-similarity SMOS of 4.0/5 — comparable to major English corpora, and openly released to accelerate Persian speech research.

ParsVoice data pipeline: audiobooks are boundary-detected, transcribed, sentence-checked, and boundary-optimized; then quality-filtered, speaker-clustered, punctuation-restored, and assembled into the final ParsVoice corpus.
The nine-stage pipeline that turns raw audiobooks into ParsVoice.
1,804h
High-quality audio
470+
Speakers
3.6/5
Naturalness (MOS)
4.0/5
Speaker similarity

Sample Demonstrations

Zero-shot clones from the fine-tuned model, next to the unseen reference voice each was cloned from.

Loading samples…