Results (CommonVoice-17 test, n=1500/lang, greedy, Whisper normalization + Arabic folding)
Table with columns: Lang, stock whisper-small, this model, metric| Lang | stock whisper-small | this model | metric |
|---|
| en | 11.7 | 12.2 | WER % |
| de | 15.4 | 15.0 | WER % |
| es | 12.1 | 9.7 | WER % |
| fr | 22.9 | 19.2 | WER % |
| ru | 17.3 | 13.6 | WER % |
| tr | 25.2 | 20.9 | WER % |
| cy | 74.2 | 60.7 | WER % |
| ar | 48.4 | 35.6 | WER % |
| th | 19.7 | 11.3 | CER % |
| zh | 34.4 | 14.9 | CER % |
| ka | 178.0 | 86.2 | WER % (unusable in both — model-class floor) |
Better than stock on 10 of 11 languages (−0.5 on en). Encoder frozen during fine-tuning (the
multilingual encoder is preserved); 15,000 steps, warmup+cosine, dev-checkpoint-selected; bf16.
Limitations
- Coverage = these 11 languages; fine-tuning erodes Whisper's other ~88 (unseen scripts most —
measured in the accompanying study). Use stock Whisper for languages far from this set.
- Georgian (ka) reported for transparency; not usable at the whisper-small tier.
- Read-speech domain (CommonVoice + FLEURS-validated); not benchmarked on far-field/telephony.
Acoustic conditions of the evaluation
Evaluated on crowdsourced consumer-microphone recordings with real environmental noise —
traffic, room reverb, variable devices — CommonVoice's native conditions, not studio audio.
The numbers above already include that heterogeneity. Not yet benchmarked: far-field, telephony
(8 kHz), overlapping speech.
(Note: this is the unconverted control — it has no MHA→MLA conversion, so "compression cost"
does not apply to it. The SNR-robustness-of-compression result lives on the MLA cards, where it
is measured as MLA-vs-this-control.)