Whisper Small Malayalam
Fine-tuned version of openai/whisper-small on a multi-corpus Malayalam speech dataset. This is the first publicly available Whisper model fine-tuned specifically for Malayalam ASR.
Model Description
- Base model: openai/whisper-small (244M parameters)
- Language: Malayalam (ml)
- Task: Automatic Speech Recognition (transcription)
- Training steps: 3500
- Best WER: 37.64% on CommonVoice 25 Malayalam test set
- Multi-source test sample: 42.9% WER / 12.3% CER (300 clips across 6 sources — see Multi-source evaluation)
- CPU speed (Transformers, FP32, 4 vCPU): RTF 1.96, i.e. slower than real time. For CPU deployment use the whisper.cpp builds
Training Data
The model was trained on an aggregated corpus of 5 Malayalam speech datasets, combined and published as sajilck/malayalam-asr-corpus.
| Corpus | Source | Domain | Access |
|---|---|---|---|
| IMaSC | thennal/imasc | TTS / Read speech | HuggingFace |
| SMC Malayalam Speech Corpus | sajilck/smc-malayalam-speech-corpus | Read speech | Kaggle |
| IndicTTS Malayalam | kavyamanohar/indic-tts-malayalam-speech-corpus | TTS / Read speech | Kaggle |
| OpenSLR 63 | sajilck/openslr63 | Crowdsourced | Kaggle |
| CommonVoice 25 Malayalam | sajilck/common-voice-malayalam | Crowdsourced | Kaggle |
Total: ~86,000 samples across TTS-recorded, read speech, and crowdsourced domains.
Benchmark Results
| Model | Params | WER ↓ | Notes |
|---|---|---|---|
| openai/whisper-small (base) | 244M | ~85% | No Malayalam fine-tuning |
| smcproject/Malwhisper-v1-medium | 769M | 61.84% | Single corpus (IMaSC only) |
| sajilck/whisper-small-malayalam | 244M | 37.64% | Multi-corpus fine-tuning |
Key advantages over prior work:
- 3× smaller model than Malwhisper-v1-medium, better WER
- 5 corpora vs 1 — better speaker and domain diversity
- Multi-domain training — TTS, read speech, and crowdsourced audio
Multi-source evaluation
The CommonVoice figure above covers one domain. To see how the model does across all domains, 300 clips (50 per source, 23.8 min of audio) were drawn at random from the corpus's test split and scored with FP32 greedy decoding.
| Metric | Result |
|---|---|
| WER, normalised (95% bootstrap CI) | 42.9% (39.5% – 46.5%) |
| WER, raw | 45.3% |
| CER, normalised | 12.3% |
| Substitutions / deletions / insertions | 30.9% / 8.1% / 3.8% |
| Empty outputs / repetition loops | 0 / 0 |
| Source | WER | CER |
|---|---|---|
| IMaSC | 23.9% | 4.0% |
| OpenSLR 63 | 37.8% | 7.5% |
| IndicTTS | 40.2% | 12.0% |
| CommonVoice | 51.4% | 11.8% |
| Shrutilipi | 52.9% | 21.7% |
| SMC | 53.0% | 11.7% |
Normalisation steps: Unicode NFC, old-style chillu (consonant + virama + ZWJ) converted to atomic chillu, zero-width joiners removed, and punctuation stripped.
How to read these numbers
- The test split overlaps with training data (see Limitations). IMaSC, SMC and OpenSLR 63 test rows are mostly duplicated in train, so those per-source scores, especially IMaSC, are probably optimistic. A re-run on the leakage-free held-out set is pending.
- CER is the fairer headline for Malayalam. CER (12.3%) is much lower than WER (42.9%). Many word "errors" are splitting or joining differences in compound words, for example മുഖ്യമന്ത്രി vs മുഖ്യ മന്ത്രി, where every character is correct.
- The 51.4% CommonVoice figure does not replace 37.64%. It comes from 50 CommonVoice-sourced rows of this corpus's own split, with a wide confidence interval and different normalisation. It is not the official CommonVoice 25 test set.
Limitations
This section is maintained openly and updated as new evaluation work uncovers issues. Last updated after the multi-source CPU evaluation (September 2026), following the earlier leakage audit of the training corpus.
Original benchmark was narrow. The initially reported 37.64% WER was measured only on the CommonVoice Malayalam test split — read, studio-quality speech. It was not evaluated against the other four source domains in the training corpus (IMaSC, SMC, IndicTTS, OpenSLR 63) or against broadcast/radio-news-style audio.
The corpus's official train/test split has confirmed leakage. An
exact-transcript-hash and fuzzy near-duplicate check between this
model's training corpus's train (86,911 rows) and test (4,828 rows)
splits found:
- 41.6% of test rows have an exact-duplicate transcript in train
- 44.4% are flagged by fuzzy near-duplicate matching
Any WER computed on the raw test split partly reflects memorization,
not generalization.
Leakage is concentrated in the smaller, fixed-script sources. Per-source kept rate after filtering out flagged rows:
| Source | Kept % |
|---|---|
| IMaSC | 2.7% |
| SMC | 5.7% |
| OpenSLR 63 | 13.9% |
| CommonVoice | 70.7% |
| IndicTTS | 84.6% |
| Shrutilipi | 93.0% |
IMaSC and SMC — small corpora with a bounded set of scripted sentences read by a handful of speakers — are almost entirely duplicated across the split. Shrutilipi, despite being scraped at document scale, is the cleanest source by this measure.
The clean held-out set is source-imbalanced. After filtering, 2,684 clean rows remain, of which 87.6% are Shrutilipi. Current held-out evaluation is better powered to measure performance on radio-news-style audio than on read/studio speech.
The GGML/GGUF "WER gap" is mostly explained by the evaluation data.
whisper.cpp f16/q5_0/q8_0 conversions scored 51.7%–54.4% WER on a
20-clip sample of the multi-source test split, far above the 37.64%
CommonVoice figure. A later CPU evaluation of the original FP32
Transformers model on the same split scored 42.9% on 300 clips and
51.1% on a 24-clip subset. The unconverted model shows a gap of the
same size, so most of it comes from the broader, harder source mix and
the noise of small samples, not from the conversion. Quantization itself
still costs a little accuracy: about 2.7 WER points from f16 to q8_0, and
about 5.6 points for PyTorch INT8 dynamic quantization on the 24-clip
subset.
What's fixed: a leakage-free held-out set now exists for evaluation.
Still open: source imbalance in that set, and re-running the
multi-source and CPU evaluation on it. The numbers above come from the
full test split.
General limitations
- Trained on read speech and crowdsourced audio, so it may perform worse on spontaneous, conversational Malayalam
- Foreign proper nouns (English names, place names) may be transcribed with Malayalam phonetic approximations
- Performance may vary across Malayalam dialects
- Not real-time on CPU with plain Transformers (see CPU inference benchmark)
Full writeup with methodology: I Checked My Own ASR Dataset for Leakage — Here's What I Found
Usage
HuggingFace Transformers (Python)
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="sajilck/whisper-small-malayalam",
generate_kwargs={"language": "malayalam", "task": "transcribe"},
)
result = pipe("your_audio.wav")
print(result["text"])
For longer audio files:
pipe = pipeline(
"automatic-speech-recognition",
model="sajilck/whisper-small-malayalam",
generate_kwargs={"language": "malayalam", "task": "transcribe"},
chunk_length_s=30,
stride_length_s=5,
)
result = pipe("long_audio.wav")
print(result["text"])
CPU Inference Benchmark (Transformers, no GPU)
This measures the model on CPU through plain HuggingFace Transformers with default settings. It is a baseline, not an optimised setup.
| Setting | Value |
|---|---|
| Hardware | Kaggle CPU: Intel Xeon @ 2.20 GHz, 4 vCPU (2 physical cores), 33.7 GB RAM |
| Software | PyTorch 2.10 (CPU), Transformers 5.0 |
| Decoding | FP32, greedy, batch size 1, max_new_tokens=225 |
| Evaluation data | 300 clips (average 4.8 s long), same sample as Multi-source evaluation |
| Metric | Result |
|---|---|
| Real-time factor (RTF, all clips together) | 1.96 (0.51× real time) |
| RTF median / p90 per clip | 2.05 / 2.97 |
| Latency per clip, p50 / p90 / max | 9.1 s / 14.1 s / 15.9 s |
| Cold start (first clip, 3.5 s long) | 4.5 s |
| Decoding speed | |
| Model load time | 13.6 s |
| Peak RAM | ~3.9 GB |
RTF = processing time ÷ audio duration. Values above 1 mean slower than real time.
Where the time goes. Feature extraction takes about 0.1% of the time; almost all of it is the decoder generating text one token at a time. Whisper's tokenizer splits Malayalam script into many small pieces: this run produced about 120 tokens per clip, roughly 25 tokens per second of audio. Speed on CPU is therefore limited mainly by transcript length in tokens, not by audio length.
| Clip length | Clips | RTF | p90 latency |
|---|---|---|---|
| < 3 s | 67 | 2.60 | 8.3 s |
| 3–6 s | 169 | 2.09 | 12.1 s |
| 6–10 s | 54 | 1.72 | 14.7 s |
| 10–15 s | 9 | 1.16 | 14.9 s |
Settings for faster inference (24-clip subset; compare these rows with each other only)
| Setting | WER | RTF | p90 latency | Size |
|---|---|---|---|---|
| FP32, greedy (default) | 51.1% | 1.80 | 14.5 s | 967 MB |
INT8 dynamic quantisation (torch.ao), greedy |
56.7% | 1.40 | 11.4 s | 413 MB |
| FP32, beam search (5 beams) | 51.7% | 4.80 | 41.1 s | 967 MB |
| CPU threads | RTF | Speed-up vs 1 thread |
|---|---|---|
| 1 | 3.18 | 1.00× |
| 2 | 1.83 | 1.73× |
| 4 | 1.81 | 1.76× |
- Use greedy decoding. Beam search with 5 beams was about 2.7× slower and no more accurate.
- Set threads to the number of physical cores. Hyperthreads added almost nothing (2 → 4 threads: +1%).
- Real-time streaming isn't feasible with this setup. With 6 s windows and a 4 s hop, p90 latency per window was 12.9 s, about 3.2× over budget. For CPU use, run the whisper.cpp builds below. The RTFs in that table were measured on 30 s audio, so they aren't directly comparable with the numbers here.
Evaluation notebook: whisper-malayalam-cpu-eval.ipynb (Kaggle, CPU only).
GGML / whisper.cpp — CPU Inference (No GPU Required)
Quantized GGML variants are available for use with whisper.cpp, enabling Malayalam ASR on any laptop or edge device without a GPU.
| Variant | File | Size | RTF (30s audio) | Speed | Notes |
|---|---|---|---|---|---|
| FP16 | ggml-model-f16.bin | 487 MB | 0.40 | 2.6× real-time | Best quality |
| Q5_0 | ggml-model-q5_0.bin | 175 MB | 0.44 | 2.3× real-time | Smallest size |
| Q8_0 | ggml-model-q8_0.bin | 264 MB | 0.34 | 3.0× real-time | ✅ Recommended |
Benchmarked on Kaggle CPU (4 cores, AVX2). Q8_0 outperforms F16 and Q5_0 due to efficient SIMD integer operations on AVX2 hardware. WER difference between F16 and Q8_0 is only 2.68%, making Q8_0 the best overall choice.
Note: RTF is high for short clips (<10s) due to fixed 30s mel spectrogram encoding overhead. All variants process 30s+ audio faster than real-time.
Quick Start
# 1. Build whisper.cpp
git clone https://github.com/ggml-org/whisper.cpp
cd whisper.cpp
cmake -B build && cmake --build build --config Release
# 2. Download Q8_0 (recommended)
wget https://huggingface.co/sajilck/whisper-small-malayalam/resolve/main/ggml/ggml-model-q8_0.bin
# 3. Transcribe Malayalam audio (16kHz mono WAV)
./build/bin/whisper-cli -m ggml-model-q8_0.bin -l ml -f your_audio.wav
Audio Requirements
- Format: WAV (16-bit PCM)
- Sample rate: 16 kHz
- Channels: Mono
Convert any audio to the required format:
ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav
Training Details
| Parameter | Value |
|---|---|
| Base model | openai/whisper-small |
| Training steps | 3500 |
| Effective batch size | 16 (batch=4, grad_accum=4) |
| Learning rate | 1e-5 |
| Warmup steps | 500 |
| Precision | fp16 |
| Hardware | NVIDIA Tesla P100 16GB |
| Framework | HuggingFace Transformers + Seq2SeqTrainer |
Training Data Preprocessing
Audio from all 5 corpora was:
- Resampled to 16kHz mono
- Filtered to 0.5–30 second clips
- Converted to log-mel spectrograms (80 mel bins)
- Tokenized using Whisper's multilingual tokenizer with language token
<|ml|>
Future Work
- v2: Adding Shrutilipi broadcast news corpus with higher learning rate (lr=2e-4) based on findings from Adalat AI Vividh-ASR paper
- whisper-medium-malayalam and whisper-tiny-malayalam variants
- Evaluation on Vividh-ASR benchmark across all 4 speech difficulty tiers
- Re-run the multi-source and CPU evaluation on the leakage-free held-out set
- Faster CPU inference for real-time streaming: benchmark CTranslate2 / faster-whisper INT8 and whisper.cpp on the same clips, targeting a 6 s window in under 4 s
Citation
@misc{sajilck2026whispermalayalam,
author = {Sajil C.K.},
title = {Whisper Small Malayalam: Multi-Corpus Fine-Tuning of Whisper for Malayalam ASR},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/sajilck/whisper-small-malayalam}
}
License
This model is released under the Apache 2.0 license, consistent with the base Whisper model. Training corpora retain their individual licenses — please refer to each source dataset for usage terms.
- Downloads last month
- 160
Model tree for sajilck/whisper-small-malayalam
Base model
openai/whisper-smallDataset used to train sajilck/whisper-small-malayalam
Space using sajilck/whisper-small-malayalam 1
Evaluation results
- WER on Common Voice 25 (Malayalam)test set self-reported37.640
- WER (normalised) on malayalam-asr-corpus test (300-clip multi-source sample, partial train overlap)test set self-reported42.850
- CER (normalised) on malayalam-asr-corpus test (300-clip multi-source sample, partial train overlap)test set self-reported12.280