Key Features
- Robustness-tuned Kazakh ASR model
- Based on Kazakh Whisper Large-v3 Turbo
- Built on Whisper Large-v3 Turbo architecture
- Fine-tuned for 2 robustness epochs
- Trained with noise, reverb, speed, and volume perturbations
- Ready-to-use Transformers checkpoint
- Compatible with Hugging Face pipelines
- Suitable for noisy Kazakh ASR experiments and evaluation
- Useful for transcription, subtitles, voice assistants, and research
Important Note
This model is not intended to fully replace the original Kazakh Whisper Large-v3 Turbo in every situation.
Compared to the original model:
- Clean FLEURS performance is slightly worse.
- Robust/noisy evaluation performance is slightly better.
- The main improvements appear under stronger additive noise conditions.
For clean high-quality audio, the original model remains a very strong choice.
For noisy or perturbed audio, this robust version may be preferable.
Benchmark Summary
Clean FLEURS Kazakh Test
Table with columns: Model, WER ↓, CER ↓| Model | WER ↓ | CER ↓ |
|---|
| Kazakh Whisper Large-v3 Turbo | 11.74% | 4.97% |
| Kazakh Whisper Large-v3 Turbo Robust 2epoch | 11.90% | 5.00% |
| Whisper Large-v3 Turbo | 21.16% | 6.12% |
| Wav2Vec2 XLSR Kazakh | 21.76% | 6.27% |
| Whisper Large-v3 | 33.10% | 7.77% |
| Whisper Medium |
Clean FLEURS performance is almost preserved, with a small WER increase compared to the original Kazakh Whisper Large-v3 Turbo model.
Robust FLEURS Kazakh
Robust FLEURS Kazakh was created from the Google FLEURS Kazakh test set by applying controlled perturbations to the original test audio.
Conditions:
- clean
- noise_10db
- noise_5db
- reverb
- speed_0.9
- speed_1.1
- volume_low
- volume_high
Robust FLEURS Results
Table with columns: Model, Clean WER ↓, Robust-only WER ↓, Robust-only CER ↓| Model | Clean WER ↓ | Robust-only WER ↓ | Robust-only CER ↓ |
|---|
| Kazakh Whisper Large-v3 Turbo | 11.76% | 17.43% | 6.84% |
| Kazakh Whisper Large-v3 Turbo Robust 2epoch | 11.94% | 17.33% | 6.74% |
The robust 2epoch model gives a small improvement on the robust-only subset:
Table with columns: Metric, Original, Robust 2epoch, Difference| Metric | Original | Robust 2epoch | Difference |
|---|
| Robust-only WER ↓ | 17.43% | 17.33% | -0.10 |
| Robust-only CER ↓ | 6.84% | 6.74% | -0.10 |
The improvement is small and mainly appears under additive noise conditions.
Noise-only FLEURS Kazakh
A separate noise-only benchmark was created from Google FLEURS Kazakh test audio.
This benchmark focuses on synthetic stationary additive noise.
Conditions:
- clean
- white_20db
- white_15db
- white_10db
- white_5db
- white_0db
- pink_20db
- pink_15db
- pink_10db
- pink_5db
- pink_0db
White noise is additive white Gaussian noise.
Pink noise is synthetic pink noise generated using approximate 1/sqrt(f) spectral shaping.
Noise-only FLEURS Results
Table with columns: Model, Clean WER ↓, Robust-only WER ↓, Robust-only CER ↓| Model | Clean WER ↓ | Robust-only WER ↓ | Robust-only CER ↓ |
|---|
| Kazakh Whisper Large-v3 Turbo | 11.76% | 25.43% | 9.88% |
| Kazakh Whisper Large-v3 Turbo Robust 2epoch | 11.94% | 25.15% | 9.69% |
| Whisper Large-v3 | 33.18% | 54.09% | 18.54% |
| Whisper Medium | 53.13% | 73.87% |
On the noise-only benchmark, the robust 2epoch model improves over the original model:
Table with columns: Metric, Original, Robust 2epoch, Difference| Metric | Original | Robust 2epoch | Difference |
|---|
| Robust-only WER ↓ | 25.43% | 25.15% | -0.28 |
| Robust-only CER ↓ | 9.88% | 9.69% | -0.19 |
Detailed Noise-only WER
Table with columns: Condition, Kazakh Whisper Large-v3 Turbo, Robust 2epoch| Condition | Kazakh Whisper Large-v3 Turbo | Robust 2epoch |
|---|
| clean | 11.76% | 11.94% |
| white_20db | 14.37% | 14.51% |
| white_15db | 17.67% | 17.38% |
| white_10db | 23.95% | 23.92% |
| white_5db | 35.33% | 34.68% |
| white_0db | 51.23% |
Usage
Installation
pip install transformers accelerate torch torchaudio soundfile
Quick Start
The model can be used directly with the Hugging Face pipeline() API:
from transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch") result = asr( "audio.wav", generate_kwargs={ "language": "kk", "task": "transcribe" }) print(result["text"])
GPU Inference
For faster inference, use GPU inference with FP16:
import torchfrom transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch", torch_dtype=torch.float16, device="cuda") result = asr( "audio.wav", generate_kwargs={ "language": "kk", "task": "transcribe" }) print(result["text"])
CPU Inference
from transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch", device="cpu") result = asr( "audio.wav", generate_kwargs={ "language": "kk", "task": "transcribe" }) print(result["text"])
Long Audio
For long audio files such as meetings, podcasts, interviews, and lectures, use chunked inference:
from transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch", chunk_length_s=30, batch_size=8) result = asr( "meeting.wav", generate_kwargs={ "language": "kk", "task": "transcribe" }) print(result["text"])
Batch Processing
from transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch") files = [ "audio1.wav", "audio2.wav", "audio3.wav"] for file in files: result = asr( file, generate_kwargs={ "language": "kk", "task": "transcribe" } ) print(file, result["text"])
The model can also be used directly with WhisperProcessor and WhisperForConditionalGeneration:
import torchimport torchaudiofrom transformers import WhisperProcessor, WhisperForConditionalGeneration model_id = "shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch" processor = WhisperProcessor.from_pretrained(model_id) model = WhisperForConditionalGeneration.from_pretrained( model_id, torch_dtype=torch.float16, device_map="auto") audio_path = "audio.wav" speech_array, sampling_rate = torchaudio.load(audio_path) if sampling_rate != 16000: resampler = torchaudio.transforms.Resample(sampling_rate, 16000) speech_array = resampler(speech_array) audio_array = speech_array.squeeze().numpy() inputs = processor( audio_array, sampling_rate=16000, return_tensors="pt") input_features = inputs.input_features.to( model.device, dtype=torch.float16) with torch.no_grad(): predicted_ids = model.generate( input_features, language="kk", task="transcribe", max_new_tokens=225 ) text = processor.batch_decode( predicted_ids, skip_special_tokens=True)[0] print("Prediction:", text)
Dataset Annotation
from transformers import pipeline asr = pipeline( "automatic-speech-recognition", model="shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch") audio_files = [ "sample1.wav", "sample2.wav", "sample3.wav"] annotations = [] for audio_file in audio_files: result = asr( audio_file, generate_kwargs={ "language": "kk", "task": "transcribe" } ) annotations.append({ "audio": audio_file, "transcript": result["text"] }) print(annotations)
Best Practices
For best performance:
- Use 16 kHz audio when possible.
- Use
language="kk" and task="transcribe".
- Segment very long recordings before transcription.
- Use chunked inference for meetings, podcasts, interviews, and lectures.
- Use GPU inference for large-scale workloads.
- Evaluate on your own target-domain audio before production deployment.
Training Data
This model was fine-tuned from shyngys879/kazakh-whisper-large-v3-turbo using robustness-augmented Kazakh speech data.
The robustness training data was based on real-speech Kazakh data selected from issai/Kazakh_Speech_Corpus_2, excluding TTS audio.
The selected clean subset contained approximately 300 hours of real Kazakh speech.
It was augmented with multiple perturbation types, producing approximately 2,400 hours of robustness-focused training audio.
Augmentation conditions included:
- clean
- additive noise
- reverberation
- speed changes
- volume changes
Intended Applications
This model can be used for:
- Speech transcription
- Subtitle generation
- Podcast transcription
- Interview transcription
- Meeting transcription
- Voice assistants
- Noisy Kazakh ASR experiments
- Dataset annotation
- Educational applications
- Call-center analytics
- Kazakh NLP pipelines
Model Details
Table with columns: Field, Value| Field | Value |
|---|
| Model Name | Kazakh Whisper Large-v3 Turbo Robust 2epoch |
| Architecture | Whisper |
| Base Model | shyngys879/kazakh-whisper-large-v3-turbo |
| Original Base Architecture | Whisper Large-v3 Turbo |
| Language | Kazakh |
| Parameters | 0.8B |
| Fine-tuning Type | Full model fine-tuning |
| Robustness Training Data | ~300h clean real speech augmented to ~2,400h |
| Effective Robustness Epochs |
Limitations
Performance may degrade on:
- Heavy real-world background noise
- Overlapping speakers
- Strong accents or dialects
- Code-switched Kazakh/Russian speech
- Music-heavy recordings
- Far-field audio
- Call-center audio with compression artifacts
- Very long audio without segmentation
Important evaluation limitation:
The noise-only benchmark uses synthetic white and pink noise. It is useful for controlled comparison, but it is not a full replacement for real-world noisy audio evaluation with babble noise, street noise, music, codec artifacts, and far-field microphones.
Citation
@misc{sovetkhan2026kazakhwhisperrobust, title={Kazakh Whisper Large-v3 Turbo Robust 2epoch}, author={Shyngys Sovetkhan}, year={2026}, howpublished={Hugging Face Model Hub}, url={https://huggingface.co/shyngys879/kazakh-whisper-large-v3-turbo-robust-2epoch}}
Original Kazakh Whisper Large-v3 Turbo:
shyngys879/kazakh-whisper-large-v3-turbo