Intended use
The model is intended for transcribing speech in the 51 languages listed above, including for applications such as media monitoring, call-centre and citizen-feedback analysis, agricultural and health advisory services, voice interfaces, and as a component in speech translation or speech-to-speech pipelines. It is also intended as a starting point for further fine-tuning on a specific language, domain or accent.
The model transcribes one language per clip: the language token is supplied at decode time and is not detected automatically. It does not translate, does not produce timestamps or speaker labels in this configuration, and is not a speaker-identification, language-identification or voice-biometric system. It should not be used as the sole basis for consequential decisions about individuals (for example in legal, clinical, immigration or employment settings) without human review of the transcript, particularly for the languages in the weaker half of the results tables.
Usage
You can use this Colab Notebook to try out the model. It lets you transcribe an audio sample from a Hugging Face dataset, an uploaded audio file, or your own recording from your computer's microphone.
The model is used in much the same way as the base Whisper model. For better accuracy, specify the language during generation.
See the example below, which transcribes an audio sample from a Hugging Face dataset:
import transformers
import datasets
import torch
SAMPLE_RATE = 16000
LANGUAGE_TOKENS_WHISPER = {
"eng": 50259, "fra": 50265, "swa": 50318, "sna": 50324, "yor": 50325, "som": 50326,
"afr": 50327, "amh": 50334, "mlg": 50349, "lin": 50353, "hau": 50354,
"ach": 50357, "aka": 50356, "bam": 50355, "bem": 50352, "ber": 50351,
"cgg": 50350, "dag": 50348, "dga": 50347, "ewe": 50346, "ful": 50345,
"ibo": 50344, "kab": 50343, "kau": 50342, "kik": 50341, "kin": 50340,
"kln": 50339, "koo": 50338, "kpo": 50337, "led": 50336, "lgg": 50335,
"lth": 50333, "lug": 50332, "luo": 50331, "luy": 50330, "myx": 50329,
"nbl": 50328, "nya": 50323, "nyn": 50322, "orm": 50321, "pcm": 50320,
"ruc": 50319, "rwm": 50317, "sot": 50316, "teo": 50315, "tsn": 50314,
"ttj": 50313, "wol": 50312, "xho": 50311, "xog": 50310, "zul": 50309
}
LANGUAGE_NAMES = {
'Acholi': 'ach', 'Afrikaans': 'afr', 'Akan': 'aka', 'Amharic': 'amh', 'Ateso': 'teo',
'Bambara': 'bam', 'Bemba': 'bem', 'Berber': 'ber', 'Chichewa': 'nya', 'Dagaare': 'dga',
'Dagbani': 'dag', 'English': 'eng', 'Ewe': 'ewe', 'French': 'fra', 'Fulani': 'ful',
'Hausa': 'hau', 'Igbo': 'ibo', 'Ikposo': 'kpo', 'Kabyle': 'kab', 'Kalenjin': 'kln',
'Kanuri': 'kau', 'Kikuyu': 'kik', 'Kinyarwanda': 'kin', 'Kwamba': 'rwm', 'Lendu': 'led',
'Lingala': 'lin', 'Luganda': 'lug', 'Lugbara': 'lgg', 'Luhya': 'luy', 'Lumasaba': 'myx',
'Luo': 'luo', 'Lusoga': 'xog', 'Malagasy': 'mlg', 'Ndebele': 'nbl', 'Nigerian Pidgin': 'pcm',
'Oromo': 'orm', 'Rukiga': 'cgg', 'Rukonjo': 'koo', 'Runyankole': 'nyn', 'Ruruuli': 'ruc',
'Rutooro': 'ttj', 'Shona': 'sna', 'Somali': 'som', 'Sotho': 'sot', 'Swahili': 'swa',
'Thur': 'lth', 'Tswana': 'tsn', 'Wolof': 'wol', 'Xhosa': 'xho', 'Yoruba': 'yor', 'Zulu': 'zul'
}
model_id = "Sunbird/SunflowerASR-51-african-languages"
model = transformers.WhisperForConditionalGeneration.from_pretrained(model_id)
processor = transformers.WhisperProcessor.from_pretrained(model_id)
def transcribe_by_whisper(audio_array, language):
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
input_features = processor(
audio_array, sampling_rate=SAMPLE_RATE, do_normalize=True, return_tensors="pt"
).input_features.to(device)
lang_tok = LANGUAGE_TOKENS_WHISPER[LANGUAGE_NAMES[language]]
transcribe_tok = processor.tokenizer.convert_tokens_to_ids("<|transcribe|>")
notimestamps_tok = processor.tokenizer.convert_tokens_to_ids("<|notimestamps|>")
forced_decoder_ids = [
(1, lang_tok),
(2, transcribe_tok),
(3, notimestamps_tok),
]
predicted_ids = model.to(device).generate(
input_features,
forced_decoder_ids=forced_decoder_ids,
num_beams=1,
do_sample=False,
)
transcription = processor.decode(
predicted_ids, skip_special_tokens=True, clean_up_tokenization_spaces=False)
print(transcription[0])
import huggingface_hub
huggingface_hub.login()
ds = datasets.load_dataset('Sunbird/salt', 'multispeaker-lug', split='test')
audio_column = 'audio'
ds = ds.cast_column(audio_column, datasets.Audio(sampling_rate=SAMPLE_RATE))
audio_array = ds[0][audio_column]['array']
lang = 'Luganda'
transcribe_by_whisper(audio_array, lang)
Two decoding notes. The language token must be forced from the table above rather than passed as a Whisper language name: the 40 languages Whisper did not originally cover are mapped onto unused token ids, so language="lug" alone will not select Luganda. And if you set max_new_tokens, keep it at 444 (Whisper's architectural ceiling of 448 decoder positions, minus the 4-token forced prefix) — Whisper's tokeniser has no Ge'ez coverage and falls back to bytes, so Amharic costs about 2.4 tokens per character against 0.4 for a median language here, and a shorter budget truncates its transcripts.
We evaluate SunflowerASR on three benchmarks: our own Sunbird Speech Benchmark, which is the only one that covers the model's full language range, and the two independent benchmarks with published baselines, AfriVox-v2 and SimbaBench. WER and CER are computed with greedy decoding, the language token forced at the decoder, and Whisper's BasicTextNormalizer applied to both hypothesis and reference. Lower is better throughout.
Sunbird Speech Benchmark
The Sunbird Speech Benchmark is assembled from the held-out test splits of the 110 dataset subsets of our training corpus, capped at 200 clips each: 20,213 clips over 51 languages drawn from 28 source corpora. It is the only one of the three benchmarks that covers every language the model supports. Note that because it is drawn from the same collections as our training data, it is in-domain for SunflowerASR and out-of-domain for the baselines; read it as a measure of coverage across 51 languages, with AfriVox-v2 and SimbaBench below providing the arm's-length comparisons.

Macro-averages over languages, comparing against Omnilingual ASR from Meta, Simba-S from UBC-NLP, Gemini 3.5 Flash from Google and GPT-4o-transcribe from OpenAI. Systems cover different subsets of the 51 languages, so the last column repeats SunflowerASR's average over exactly the languages each baseline scores:
Table with columns: System, Languages scored, WER, CER, SunflowerASR on the same languages (WER / CER)| System | Languages scored | WER | CER | SunflowerASR on the same languages (WER / CER) |
|---|
| SunflowerASR (ours) | 51 | 0.293 | 0.091 | — |
| OmniASR-LLM-7B | 46 | 0.353 | 0.111 | 0.300 / 0.093 |
| OmniASR-LLM-1B | 46 | 0.383 | 0.119 | 0.300 / 0.093 |
SunflowerASR has the lowest WER of any system compared on 30 of the 51 languages, and the lowest CER on 34.
Per-language results, sorted by WER, alongside the hours of supervised training audio that language contributed to the corpus:
Table with columns: Code, Language, Training hours, Test clips, WER, CER| Code | Language | Training hours | Test clips | WER | CER |
|---|
fra | French | 11.4 | 200 | 0.079 | 0.035 |
tsn | Tswana | 235.3 | 369 | 0.093 | 0.031 |
|

AfriVox-v2
AfriVox-v2 is an in-the-wild benchmark of 251 hours aggregating the Africa Next Voices, Waxal and AfriVox v1 corpora over a 10-domain taxonomy spanning health, finance, government and agriculture. It is largely spontaneous rather than read speech. We score 111,051 clips in the 14 languages that have metrics reported in the AfriVox-v2 paper; baseline figures are taken from that paper, which reports WER only. Simba-S is our own evaluation, and does not cover three of the languages, so its average is over the other eleven. Best per row in bold.
Table with columns: Language, SunflowerASR (ours), Sahara v2, Gemini 3 Flash, Omni-CTC 7B, Omni-CTC 1B, Omni-CTC 300M, Simba-S| Language | SunflowerASR (ours) | Sahara v2 | Gemini 3 Flash | Omni-CTC 7B | Omni-CTC 1B | Omni-CTC 300M | Simba-S |
|---|
| Swahili | 0.108 | 0.071 | 0.076 | 0.077 | 0.097 | 0.152 | 0.150 |
| Tswana | 0.116 | 0.140 |
SimbaBench
SimbaBench aggregates read and conversational speech from Common Voice, BembaSpeech, Lwazi, NCHLT and regional collections. The table covers the 19 languages in common with SunflowerASR; baseline figures come from the public leaderboard. Best per row in bold.
Table with columns: Language, SunflowerASR (ours), OmniASR-LLM-7B, Simba-S, Simba-M, Gemini 2.5 Pro| Language | SunflowerASR (ours) | OmniASR-LLM-7B | Simba-S | Simba-M | Gemini 2.5 Pro |
|---|
| Hausa | 0.167 | 0.169 | 0.641 | 0.291 | 0.219 |
| Luganda | 0.220 | 0.171 | 0.227 | 0.357 | 0.421 |
| Swahili | 0.246 |
Training Datasets
The model was fine-tuned on a multilingual corpus covering 51 African languages, assembled from a range of publicly available and community-collected speech datasets. The table below lists each source dataset and the languages it contributes to training (ISO 639-3 code in parentheses). A language may appear in more than one source.
Corpus statistics
Table | |
|---|
| Languages | 51 |
| Dataset subsets | 113 (29 source collections, a mean of 2.2 independent sources per language) |
| Clips before filtering | 4.15 M |
| Removed: suspected label noise (CER > 0.4 against an intermediate checkpoint) | 4.1% |
| Removed: clips longer than Whisper's 30 s window | 3.9% |
| Clips after filtering | 3.82 M |
| Total audio after filtering | 7,412 h |
| Mean clip duration | 7.0 s |
| Packed training samples (clips concatenated into 30 s windows) |
Per-language hours are listed alongside the benchmark results in the Sunbird Speech Benchmark table above. Clips are filtered and then packed: short clips from the same subset are concatenated to fill the 30 s input window, which keeps every pack monolingual and single-source while cutting padding waste from 76.7% to 19.4%. Of the 113 training subsets, the 110 with held-out test splits make up the Sunbird Speech Benchmark.
Sources
Table with columns: Source dataset, Languages| Source dataset | Languages |
|---|
| Mozilla Common Voice | Afrikaans (afr), Amharic (amh), Rukiga (cgg), Dagbani (dag), Hausa (hau), Igbo (ibo), Kabyle (kab), Kinyarwanda (kin), Kalenjin (kln), Rukonjo (koo), Lendu (led), Thur (lth), Luganda (lug), Luo (luo), Nigerian Pidgin (pcm), Ruruuli (ruc), Kwamba (rwm), Swahili (swa), Tswana (tsn), Rutooro (ttj), Yoruba (yor) |
| Google FLEURS | Afrikaans (afr), Fulani (ful), Hausa (hau), Igbo (ibo), Lingala (lin), Luganda (lug), Luo (luo), Chichewa (nya), Oromo (orm), Shona (sna), Somali (som), Sotho (sot), Swahili (swa), Wolof (wol), Xhosa (xho), Yoruba (yor), Zulu (zul) |
| Google Waxal | Acholi (ach), Akan (aka), Amharic (amh), Dagbani (dag), Dagaare (dga), Ewe (ewe), Fulani (ful), Ikposo (kpo), Lingala (lin), Luganda (lug), Malagasy (mlg), Lumasaba (myx), Runyankole (nyn), Oromo (orm), Shona (sna), Lusoga (xog) |
| ASR Africa Data Efficiency Benchmark |
Notes
- African Next Voices is drawn from two collection hubs: Kenya (Kikuyu, Kalenjin, Luo,
Somali) and Southern Africa (Ndebele, Sotho, Tswana, Xhosa, Zulu).
Training procedure
SunflowerASR is trained from Whisper large-v3 in two stages. In the first, all 1.55 B parameters are fine-tuned for 5 epochs over the packed corpus described above, with a per-device batch of 64 packed samples and 4 gradient accumulation steps — an effective batch of 256 packed samples, or roughly 1.7 hours of audio per optimiser step, giving about 4,300 optimiser steps per epoch — at a peak learning rate of 2×10⁻⁵ with 100 warmup steps and cosine decay, label smoothing of 0.1, and augmentation applied on the fly: additive noise drawn from a corpus of Ugandan ambient recordings (Sunbird/urban-noise-uganda-61k) plus speed and bandwidth perturbation. The second stage is a reinforcement learning pass with Group Relative Policy Optimisation (GRPO), which optimises the sequence-level metric directly rather than the per-token likelihood of a single reference: for each clip the policy samples 10 complete transcripts, each is scored with the reward -min(CER, 1.0) against the reference, and the group's own mean and standard deviation turn those scores into advantages, so probability mass moves towards the better rollouts under a KL penalty of β = 0.01 against the initial policy. It runs for one epoch over 15,285 clips, selected as the 200 clips per dataset subset with the largest headroom — the gap between the median and the best CER over 10 sampled rollouts, which identifies clips where a better transcript is already inside the model's sampling distribution and only needs to be made more probable — with 64 distinct clips (640 rollouts) per update at a peak learning rate of 10⁻⁵, the encoder frozen so that only the decoder's 906 M parameters are trained. This stage leaves clean-speech accuracy roughly unchanged but improves noisy and spontaneous audio: it largely suppresses the repetition loops Whisper falls into under acoustic stress, makes ordinary substitutions closer to the reference spelling, and helps most for the languages with the least training data.
Training configuration
Table | |
|---|
| Base model | openai/whisper-large-v3 (1.55 B parameters) |
| Stage 1 | Supervised fine-tuning, all parameters, 5 epochs, LR 2×10⁻⁵ cosine, label smoothing 0.1, effective batch 256 packs |
| Stage 2 | GRPO (trl, loss_type: dapo), decoder only (906 M of 1.55 B), 1 epoch, LR 10⁻⁵ |
| GRPO reward | -min(CER, 1.0) against a single reference transcript |
| GRPO rollouts | 10 per clip, temperature 0.8, top-p 0.98; 64 clips (640 rollouts) per update |
Limitations
- Quality is very uneven across the 51 languages. On the Sunbird Speech Benchmark, WER ranges from 0.08 to 0.67. Check the per-language figures before deploying for any one language, and treat the tail of the table as usable for triage or search rather than for verbatim transcription.
- The language must be specified at decode time. The model does not perform language identification, and giving it the wrong language token degrades output substantially. Language tokens for the 40 languages Whisper did not originally support are overwritten unused tokens, so the mapping in the Usage section must be used rather than Whisper's own language names.
- Audio longer than 30 seconds must be segmented before transcription, as with any Whisper model; longer clips are truncated by the encoder.
- Orthographic convention rather than intelligibility. Both training stages score hypotheses against a single reference transcript, so the model is optimised towards the transcription conventions of these corpora. Languages without a settled orthography inherit that ambiguity, and a transcript may be penalised — or produced — in a spelling a given community would not use.
- Domain and style. Much of the training data for the lower-resource languages is read speech recorded in quiet conditions. Accuracy is lower on spontaneous, overlapping, noisy or heavily code-switched speech, though this is the setting the GRPO stage improves most.
- Repetition loops are reduced, not eliminated. Under heavy acoustic stress the model can still emit a repeated run or a hypothesis much longer than the utterance.
- Benchmark caveat. The Sunbird Speech Benchmark is drawn from the same collections as the training corpus, so it is in-domain for this model and out-of-domain for the baselines it is compared against. AfriVox-v2 and SimbaBench are the arm's-length comparisons.
Citation
If you use this model, please cite the accompanying paper:
@misc{sunflowerasr2026,
title = {Group Relative Policy Optimisation Improves Multilingual Speech
Recognition in Low-Resource Languages},
author = {Hu, Tim Wenjie and Ouma, Evelyn Nafula and Akera, Benjamin and
Quinn, John A.},
year = {2026},
institution = {Sunbird AI},
howpublished = {\url{https://huggingface.co/Sunbird/SunflowerASR-51-african-languages}}
}
And, for the evaluation set:
@misc{sunbird_speech_benchmark_2026,
title = {Sunbird Speech Benchmark: a multilingual ASR test set for 51
African languages},
author = {Sunbird AI},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/Sunbird/speech-benchmark}}
}
License and acknowledgements
Released under the Apache 2.0 licence, following the licence of the Whisper large-v3 base model. The training corpus is assembled from the publicly available and community-collected datasets credited in Training Datasets; please also respect the licence and citation requirements of those sources. We thank the many organisations and community contributors who collected and released this speech data, without which a model covering these languages would not be possible.