State-of-the-Art Persian Automatic Speech Recognition

This project showcases the development of the best-performing Automatic Speech Recognition (ASR) models for the Persian language.

Key achievements include:

  • Training top-performing models like Whisper, NeMo Parakeet, and NeMo FastConformer.
  • Achieving a state-of-the-art Word Error Rate (WER) of 4.59%, a 78% relative improvement over baselines.
  • Curating a 10,000-hour proprietary dataset, one of the largest for the Persian language.
  • Building scalable data processing pipelines for creating high-quality speech datasets from raw audio.

The models significantly outperform previous academic and proprietary systems. A live demonstration is available on Hugging Face Spaces.