Key capabilities
- local CPU or GPU inference
- individual file and folder transcription through the companion toolkit
- browser demo for recording or uploading audio
- developer-ready Transformers integration
- Somali-focused speech-to-text output
Published Somali results
- validation WER:
0.2166
- validation CER:
0.1054
- test WER:
0.2278
- test CER:
0.1186
Training overview
- base model:
openai/whisper-large-v3
- architecture:
WhisperForConditionalGeneration
- training rows:
39,604
- validation rows:
429
- test rows:
428
- training audio: approximately
72.1 hours
- epochs:
6
- learning rate:
0.0001
- primary language: Somali (
so)
A private Ogaal Labs collection contributed roughly 5,000 curated prompts recorded by 19 speakers across varied genders, accents, and speaking styles. English was not part of the training objective.
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="Ogaal-Labs/Ogaal-ASR",
)
result = pipe("test.wav")
print(result["text"])
Use the companion GitHub repository for local CPU/GPU inference and the browser demo:
git clone https://github.com/Ogaal-Labs/Ogaal-ASR.git
cd Ogaal-ASR
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
git clone https://huggingface.co/Ogaal-Labs/Ogaal-ASR model
python scripts/infer_somali_asr.py \
--audio-path /path/to/audio.wav \
--model-dir model
Requirements: Python 3.10 or newer and ffmpeg on the system path.
Files included
model.safetensors: fine-tuned Whisper weights
config.json, generation_config.json, processor_config.json: runtime configuration
tokenizer.json, tokenizer_config.json: tokenizer assets
best_val_metrics.json, test_metrics.json: held-out evaluation summaries
TECHNICAL_BOOK.md, MODEL_SCOPE.md: supporting release documentation
Intended use
- Somali speech transcription
- local and privacy-sensitive transcription workflows
- developer integration for Somali voice products
- evaluation and benchmarking on Somali audio
Limitations
- designed for Somali, not general multilingual transcription
- English was not part of the training objective
- accuracy may vary with accents, noise, microphones, domains, and speaking styles
- users should independently evaluate the model before high-stakes deployment
Links
License
This model repository is released under the Apache License 2.0. The base openai/whisper-large-v3 model is also distributed under Apache-2.0.