Lingala ASR — mms-1b-all full fine-tune (no pseudo-labels)

Component of the 4th-place solution to the Google WAXAL ASR Challenge (Zindi, phase 2): 892 unseen clips, two African languages, no language metadata, scored 1 - (WER + CER) / 2 on raw text. Private leaderboard 0.771848284.

Code, full method and one-command verification: yehoshua0/waxal-asr-phase2 The repository reproduces the submitted CSV byte for byte on a laptop in about a minute, and re-decodes every input from the audio on rented GPUs in about four hours.

Role in the system

Ablation. The primary without the pseudo-label and per-speaker stages.

What it measured

The fully reproducible chain built on this checkpoint reaches 0.747925 against the shipped 0.762400: pseudo-labels plus speaker adaptation are 93% of that gap.

Usage

from transformers import Wav2Vec2ForCTC, AutoProcessor
import torch, soundfile as sf, torchaudio.functional as AF

proc = AutoProcessor.from_pretrained("yehoshua01/waxal-mms-1b-lin-full")
model = Wav2Vec2ForCTC.from_pretrained("yehoshua01/waxal-mms-1b-lin-full").eval().cuda()

wav, sr = sf.read("clip.wav", dtype="float32")
wav = AF.resample(torch.from_numpy(wav), sr, 16_000).numpy()

f = proc(wav, sampling_rate=16_000, return_tensors="pt", padding=True)
# mms-1b-all is feat_extract_norm="layer": the attention mask is REQUIRED on batched input
logits = model(f.input_values.cuda(), attention_mask=f.attention_mask.cuda()).logits
print(proc.decode(logits[0].argmax(-1).cpu().numpy()))

The rest of the system

code, method, verification yehoshua0/waxal-asr-phase2
cached decodes and chain inputs yehoshua01/waxal-phase2-chain-inputs
all checkpoints yehoshua01 on the Hub

Sibling checkpoints (primaries, voters and ablations of the same system): waxal-mms-1b-lin-pl2-spk · waxal-sunbird51-sna-pl2-spk · waxal-whisper-turbo-lin-r1 · waxal-whisper-turbo-lin-r2 · waxal-qlora-largev3-lin · waxal-omni-ctc1b-lin · waxal-omni-ctc1b-sna · waxal-sunbird51-lin-ft-r2 · waxal-sunbird51-lin-ft-light · waxal-mms-1b-lin-fullmeta · waxal-ssa-hubert-lin

Licence and intended use

cc-by-nc-4.0. Training data is google/WaxalNLP (CC-BY-SA-4.0, share-alike), so derivatives carry that too.

These weights are not a general-purpose ASR model. Pseudo-labels were computed on the phase-2 test audio (transductive self-training, permitted for phase-2 training by the host), so the checkpoint is partly adapted to that specific set.

Downloads last month
6
Safetensors
Model size
1.0B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yehoshua01/waxal-mms-1b-lin-full

Finetuned
(457)
this model