Usage
Requires 16 kHz mono audio (the pipeline below resamples for you).
from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="theshivam7/whisper-medium-indian-english")result = pipe("audio.wav")print(result["text"])
For audio longer than 30 seconds, enable chunking:
result = pipe("long_audio.wav", chunk_length_s=30, stride_length_s=5)
Model description
A full fine-tune (all 769M parameters) of Whisper Medium via transformers Seq2SeqTrainer,
adapted to Indian English academic/lecture speech. Standard Whisper encoder-decoder architecture,
no structural changes. Input: 16 kHz mono audio; output: English text.
Intended uses & limitations
Evaluated on Indian-accented English academic lecture speech (TIE_shorts). Not evaluated on
other English varieties or domains.
⚠️ Speaker overlap (disclosed). The training data's official split shares speakers with
the test split (100% of test speakers, and 100% of test clips, are from speakers also seen in
training). There is no clip-level leakage, but part of the measured gain over the pretrained
baseline likely reflects speaker adaptation rather than purely accent/content adaptation. See
the benchmark repo for full methodology.
⚠️ No net gain, and this is the largest of a 3-size family. The +0.20pp WER change over the
pretrained baseline (below) is not statistically significant (95% CI [−0.46, +1.03] pp,
pHolm=0.774). Across the tiny/small/medium capacity study, gains shrink as model
size grows and flip to null at this size; see the benchmark repo
for the full capacity-study writeup. Do not assume this checkpoint outperforms the base model.
⚠️ Speaker-disjoint re-evaluation shows a regression, not just a null. Under a strict
speaker-disjoint re-split (train speakers never appear in test) run with 3 seeds, one seed
(42) shows a statistically significant WER regression (+1.75pp, 95% CI [+0.13, +4.17],
pHolm=0.048) against this same 14.42% pretrained baseline; the other two seeds move
the same direction (+0.38pp and +0.79pp) without reaching significance. A size-matched,
speaker-overlapping control at the same ~567-clip training budget lands flat across all three
seeds (-0.09, -0.02, -0.02 pp, all pHolm=1.000), so the regression traces to
speaker-disjointness itself, not the smaller training set.
The implication is that the +0.20pp null above is not neutral. Once speaker overlap is
removed, this fine-tune is worse than not fine-tuning at all, so the official-split number was
propped up by adaptation to test speakers. Full tables:
results/tie/analysis/finetune_disjoint_control.md.
Note that the per-clip transcripts for those six control runs were not retained, so that table
is the scoring output as recorded rather than something recomputable from the repo.
Training and evaluation data
- Train: TIE_shorts official train
split, 7,200 filtered clips (46.9h, filtered from 7,884 raw for ≤30s duration and non-empty
transcripts).
- Validation (checkpoint selection only): official validation split, 986 clips.
- Test (final evaluation): official test split, 986 clips (985 scored, one excluded for an
empty reference).
- Targets: the dataset's gold
Transcript field (English, transcribe task).
Training procedure
Training hyperparameters
- learning_rate: 1e-05
- train_batch_size: 8
- eval_batch_size: 8
- seed: 42
- gradient_accumulation_steps: 2
- total_train_batch_size: 16
- optimizer: AdamW (betas=(0.9, 0.999), epsilon=1e-08)
- lr_scheduler_type: linear
- lr_scheduler_warmup_ratio: 0.1
- num_epochs: 10 (with early stopping, patience 2; best checkpoint restored via
load_best_model_at_end)
- weight_decay: 0.01
- max_grad_norm: 1.0
- generation_max_length: 225
- regularization: SpecAugment (
mask_time_prob=0.05, mask_feature_prob=0.05)
Checkpoint selection used validation WER (same normalization as the final metric). Training
early-stopped at epoch 3 (patience 2), final training loss ≈ 0.39.
Evaluation results
Corpus WER on TIE_shorts test split, transcript_clean normalization, both models decoded
through the identical HF transformers pipeline (engine-controlled comparison):
Table with columns: Model, Corpus WER| Model | Corpus WER |
|---|
| openai/whisper-medium (pretrained baseline) | 14.42% |
| theshivam7/whisper-medium-indian-english (this model) | 14.61% |
Δ = +0.20 pp (95% CI [−0.46, +1.03], pHolm=0.774, not significant, see limitations
above).
Framework versions
- Transformers 4.46.3
- PyTorch 2.5.1
- Datasets 4.8.5
- Accelerate 1.1.1
Citation
Part of Indian-ASR-Bench. See the repo for
full benchmark methodology, error analysis, and reproduction instructions.