State-of-the-Art Persian Automatic Speech Recognition
Key achievements include:
- Training top-performing models like Whisper, NeMo Parakeet, and NeMo FastConformer.
- Achieving a state-of-the-art Word Error Rate (WER) of 4.59%, a 78% relative improvement over baselines.
- Curating a 10,000-hour proprietary dataset, one of the largest for the Persian language.
- Building scalable data processing pipelines for creating high-quality speech datasets from raw audio.
The models significantly outperform previous academic and proprietary systems. A live demonstration is available on Hugging Face Spaces.
