Whisper Small Malayalam

Fine-tuned version of openai/whisper-small on a multi-corpus Malayalam speech dataset. This is the first publicly available Whisper model fine-tuned specifically for Malayalam ASR.

Model Description

  • Base model: openai/whisper-small (244M parameters)
  • Language: Malayalam (ml)
  • Task: Automatic Speech Recognition (transcription)
  • Training steps: 3500
  • Best WER: 37.64% on CommonVoice 25 Malayalam test set
  • Multi-source test sample: 42.9% WER / 12.3% CER (300 clips across 6 sources — see Multi-source evaluation)
  • CPU speed (Transformers, FP32, 4 vCPU): RTF 1.96, i.e. slower than real time. For CPU deployment use the whisper.cpp builds

Training Data

The model was trained on an aggregated corpus of 5 Malayalam speech datasets, combined and published as sajilck/malayalam-asr-corpus.

Corpus Source Domain Access
IMaSC thennal/imasc TTS / Read speech HuggingFace
SMC Malayalam Speech Corpus sajilck/smc-malayalam-speech-corpus Read speech Kaggle
IndicTTS Malayalam kavyamanohar/indic-tts-malayalam-speech-corpus TTS / Read speech Kaggle
OpenSLR 63 sajilck/openslr63 Crowdsourced Kaggle
CommonVoice 25 Malayalam sajilck/common-voice-malayalam Crowdsourced Kaggle

Total: ~86,000 samples across TTS-recorded, read speech, and crowdsourced domains.

Benchmark Results

Model Params WER ↓ Notes
openai/whisper-small (base) 244M ~85% No Malayalam fine-tuning
smcproject/Malwhisper-v1-medium 769M 61.84% Single corpus (IMaSC only)
sajilck/whisper-small-malayalam 244M 37.64% Multi-corpus fine-tuning

Key advantages over prior work:

  • 3× smaller model than Malwhisper-v1-medium, better WER
  • 5 corpora vs 1 — better speaker and domain diversity
  • Multi-domain training — TTS, read speech, and crowdsourced audio

Multi-source evaluation

The CommonVoice figure above covers one domain. To see how the model does across all domains, 300 clips (50 per source, 23.8 min of audio) were drawn at random from the corpus's test split and scored with FP32 greedy decoding.

Metric Result
WER, normalised (95% bootstrap CI) 42.9% (39.5% – 46.5%)
WER, raw 45.3%
CER, normalised 12.3%
Substitutions / deletions / insertions 30.9% / 8.1% / 3.8%
Empty outputs / repetition loops 0 / 0
Source WER CER
IMaSC 23.9% 4.0%
OpenSLR 63 37.8% 7.5%
IndicTTS 40.2% 12.0%
CommonVoice 51.4% 11.8%
Shrutilipi 52.9% 21.7%
SMC 53.0% 11.7%

Normalisation steps: Unicode NFC, old-style chillu (consonant + virama + ZWJ) converted to atomic chillu, zero-width joiners removed, and punctuation stripped.

How to read these numbers

  • The test split overlaps with training data (see Limitations). IMaSC, SMC and OpenSLR 63 test rows are mostly duplicated in train, so those per-source scores, especially IMaSC, are probably optimistic. A re-run on the leakage-free held-out set is pending.
  • CER is the fairer headline for Malayalam. CER (12.3%) is much lower than WER (42.9%). Many word "errors" are splitting or joining differences in compound words, for example മുഖ്യമന്ത്രി vs മുഖ്യ മന്ത്രി, where every character is correct.
  • The 51.4% CommonVoice figure does not replace 37.64%. It comes from 50 CommonVoice-sourced rows of this corpus's own split, with a wide confidence interval and different normalisation. It is not the official CommonVoice 25 test set.

Limitations

This section is maintained openly and updated as new evaluation work uncovers issues. Last updated after the multi-source CPU evaluation (September 2026), following the earlier leakage audit of the training corpus.

Original benchmark was narrow. The initially reported 37.64% WER was measured only on the CommonVoice Malayalam test split — read, studio-quality speech. It was not evaluated against the other four source domains in the training corpus (IMaSC, SMC, IndicTTS, OpenSLR 63) or against broadcast/radio-news-style audio.

The corpus's official train/test split has confirmed leakage. An exact-transcript-hash and fuzzy near-duplicate check between this model's training corpus's train (86,911 rows) and test (4,828 rows) splits found:

  • 41.6% of test rows have an exact-duplicate transcript in train
  • 44.4% are flagged by fuzzy near-duplicate matching

Any WER computed on the raw test split partly reflects memorization, not generalization.

Leakage is concentrated in the smaller, fixed-script sources. Per-source kept rate after filtering out flagged rows:

Source Kept %
IMaSC 2.7%
SMC 5.7%
OpenSLR 63 13.9%
CommonVoice 70.7%
IndicTTS 84.6%
Shrutilipi 93.0%

IMaSC and SMC — small corpora with a bounded set of scripted sentences read by a handful of speakers — are almost entirely duplicated across the split. Shrutilipi, despite being scraped at document scale, is the cleanest source by this measure.

The clean held-out set is source-imbalanced. After filtering, 2,684 clean rows remain, of which 87.6% are Shrutilipi. Current held-out evaluation is better powered to measure performance on radio-news-style audio than on read/studio speech.

The GGML/GGUF "WER gap" is mostly explained by the evaluation data. whisper.cpp f16/q5_0/q8_0 conversions scored 51.7%–54.4% WER on a 20-clip sample of the multi-source test split, far above the 37.64% CommonVoice figure. A later CPU evaluation of the original FP32 Transformers model on the same split scored 42.9% on 300 clips and 51.1% on a 24-clip subset. The unconverted model shows a gap of the same size, so most of it comes from the broader, harder source mix and the noise of small samples, not from the conversion. Quantization itself still costs a little accuracy: about 2.7 WER points from f16 to q8_0, and about 5.6 points for PyTorch INT8 dynamic quantization on the 24-clip subset.

What's fixed: a leakage-free held-out set now exists for evaluation. Still open: source imbalance in that set, and re-running the multi-source and CPU evaluation on it. The numbers above come from the full test split.

General limitations

  • Trained on read speech and crowdsourced audio, so it may perform worse on spontaneous, conversational Malayalam
  • Foreign proper nouns (English names, place names) may be transcribed with Malayalam phonetic approximations
  • Performance may vary across Malayalam dialects
  • Not real-time on CPU with plain Transformers (see CPU inference benchmark)

Full writeup with methodology: I Checked My Own ASR Dataset for Leakage — Here's What I Found

Usage

HuggingFace Transformers (Python)

from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="sajilck/whisper-small-malayalam",
    generate_kwargs={"language": "malayalam", "task": "transcribe"},
)

result = pipe("your_audio.wav")
print(result["text"])

For longer audio files:

pipe = pipeline(
    "automatic-speech-recognition",
    model="sajilck/whisper-small-malayalam",
    generate_kwargs={"language": "malayalam", "task": "transcribe"},
    chunk_length_s=30,
    stride_length_s=5,
)
result = pipe("long_audio.wav")
print(result["text"])

CPU Inference Benchmark (Transformers, no GPU)

This measures the model on CPU through plain HuggingFace Transformers with default settings. It is a baseline, not an optimised setup.

Setting Value
Hardware Kaggle CPU: Intel Xeon @ 2.20 GHz, 4 vCPU (2 physical cores), 33.7 GB RAM
Software PyTorch 2.10 (CPU), Transformers 5.0
Decoding FP32, greedy, batch size 1, max_new_tokens=225
Evaluation data 300 clips (average 4.8 s long), same sample as Multi-source evaluation
Metric Result
Real-time factor (RTF, all clips together) 1.96 (0.51× real time)
RTF median / p90 per clip 2.05 / 2.97
Latency per clip, p50 / p90 / max 9.1 s / 14.1 s / 15.9 s
Cold start (first clip, 3.5 s long) 4.5 s
Decoding speed 13 tokens/s (77 ms per token)
Model load time 13.6 s
Peak RAM ~3.9 GB

RTF = processing time ÷ audio duration. Values above 1 mean slower than real time.

Where the time goes. Feature extraction takes about 0.1% of the time; almost all of it is the decoder generating text one token at a time. Whisper's tokenizer splits Malayalam script into many small pieces: this run produced about 120 tokens per clip, roughly 25 tokens per second of audio. Speed on CPU is therefore limited mainly by transcript length in tokens, not by audio length.

Clip length Clips RTF p90 latency
< 3 s 67 2.60 8.3 s
3–6 s 169 2.09 12.1 s
6–10 s 54 1.72 14.7 s
10–15 s 9 1.16 14.9 s

Settings for faster inference (24-clip subset; compare these rows with each other only)

Setting WER RTF p90 latency Size
FP32, greedy (default) 51.1% 1.80 14.5 s 967 MB
INT8 dynamic quantisation (torch.ao), greedy 56.7% 1.40 11.4 s 413 MB
FP32, beam search (5 beams) 51.7% 4.80 41.1 s 967 MB
CPU threads RTF Speed-up vs 1 thread
1 3.18 1.00×
2 1.83 1.73×
4 1.81 1.76×
  • Use greedy decoding. Beam search with 5 beams was about 2.7× slower and no more accurate.
  • Set threads to the number of physical cores. Hyperthreads added almost nothing (2 → 4 threads: +1%).
  • Real-time streaming isn't feasible with this setup. With 6 s windows and a 4 s hop, p90 latency per window was 12.9 s, about 3.2× over budget. For CPU use, run the whisper.cpp builds below. The RTFs in that table were measured on 30 s audio, so they aren't directly comparable with the numbers here.

Evaluation notebook: whisper-malayalam-cpu-eval.ipynb (Kaggle, CPU only).

GGML / whisper.cpp — CPU Inference (No GPU Required)

Quantized GGML variants are available for use with whisper.cpp, enabling Malayalam ASR on any laptop or edge device without a GPU.

Variant File Size RTF (30s audio) Speed Notes
FP16 ggml-model-f16.bin 487 MB 0.40 2.6× real-time Best quality
Q5_0 ggml-model-q5_0.bin 175 MB 0.44 2.3× real-time Smallest size
Q8_0 ggml-model-q8_0.bin 264 MB 0.34 3.0× real-time ✅ Recommended

Benchmarked on Kaggle CPU (4 cores, AVX2). Q8_0 outperforms F16 and Q5_0 due to efficient SIMD integer operations on AVX2 hardware. WER difference between F16 and Q8_0 is only 2.68%, making Q8_0 the best overall choice.

Note: RTF is high for short clips (<10s) due to fixed 30s mel spectrogram encoding overhead. All variants process 30s+ audio faster than real-time.

Quick Start

# 1. Build whisper.cpp
git clone https://github.com/ggml-org/whisper.cpp
cd whisper.cpp
cmake -B build && cmake --build build --config Release

# 2. Download Q8_0 (recommended)
wget https://huggingface.co/sajilck/whisper-small-malayalam/resolve/main/ggml/ggml-model-q8_0.bin

# 3. Transcribe Malayalam audio (16kHz mono WAV)
./build/bin/whisper-cli -m ggml-model-q8_0.bin -l ml -f your_audio.wav

Audio Requirements

  • Format: WAV (16-bit PCM)
  • Sample rate: 16 kHz
  • Channels: Mono

Convert any audio to the required format:

ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav

Training Details

Parameter Value
Base model openai/whisper-small
Training steps 3500
Effective batch size 16 (batch=4, grad_accum=4)
Learning rate 1e-5
Warmup steps 500
Precision fp16
Hardware NVIDIA Tesla P100 16GB
Framework HuggingFace Transformers + Seq2SeqTrainer

Training Data Preprocessing

Audio from all 5 corpora was:

  • Resampled to 16kHz mono
  • Filtered to 0.5–30 second clips
  • Converted to log-mel spectrograms (80 mel bins)
  • Tokenized using Whisper's multilingual tokenizer with language token <|ml|>

Future Work

  • v2: Adding Shrutilipi broadcast news corpus with higher learning rate (lr=2e-4) based on findings from Adalat AI Vividh-ASR paper
  • whisper-medium-malayalam and whisper-tiny-malayalam variants
  • Evaluation on Vividh-ASR benchmark across all 4 speech difficulty tiers
  • Re-run the multi-source and CPU evaluation on the leakage-free held-out set
  • Faster CPU inference for real-time streaming: benchmark CTranslate2 / faster-whisper INT8 and whisper.cpp on the same clips, targeting a 6 s window in under 4 s

Citation

@misc{sajilck2026whispermalayalam,
  author = {Sajil C.K.},
  title = {Whisper Small Malayalam: Multi-Corpus Fine-Tuning of Whisper for Malayalam ASR},
  year = {2026},
  publisher = {HuggingFace},
  url = {https://huggingface.co/sajilck/whisper-small-malayalam}
}

License

This model is released under the Apache 2.0 license, consistent with the base Whisper model. Training corpora retain their individual licenses — please refer to each source dataset for usage terms.

Downloads last month
160
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sajilck/whisper-small-malayalam

Finetuned
(3792)
this model

Dataset used to train sajilck/whisper-small-malayalam

Space using sajilck/whisper-small-malayalam 1

Evaluation results

  • WER on Common Voice 25 (Malayalam)
    test set self-reported
    37.640
  • WER (normalised) on malayalam-asr-corpus test (300-clip multi-source sample, partial train overlap)
    test set self-reported
    42.850
  • CER (normalised) on malayalam-asr-corpus test (300-clip multi-source sample, partial train overlap)
    test set self-reported
    12.280