Base model
Qwen/Qwen3-ASR-1.7B
Training data
Training used a mixture of Persian public/crawled speech and elderly-speech resources, including Common Voice Persian, Ganjoor-derived data, Filimo-derived data, locally collected elderly Persian speech, and cross-lingual elderly speech data. Audio normalization and probabilistic augmentation such as noise injection, speed perturbation, signal degradation, and pause insertion were used.
Evaluation
The reported aggregate result for the fine-tuned Qwen system is WER 0.31 and CER 0.27. Results depend on the exact normalization and test composition; model comparisons and evaluation code will be published separately.
Intended use and limitations
Intended for Persian ASR research and prototyping. Performance may vary across accents, recording devices, noise conditions, ages, and domains. Outputs can contain omissions or substitutions and should not be treated as authoritative transcripts in safety-critical settings.
Authors
Ali Alvandi — undergraduate project supervised by Hossein Sameti.