KASA-42 (Kusaal, third-party export)

Author: Prince Nasamu Alhassan

Overview

Not a model this project trained. It is the Kusaal recogniser that held the bake-off crown at 41.0 WER, republished here so the bench could score it.

It cannot be fine-tuned from. It ships raw .pt files and an inference-only int8 ONNX export with no architectures and no model_type, so nothing can start from it. That is the whole reason tekyerema-asr-mms-kus exists: it scores 30.44, beats this by 10.6 points, and can be trained further.

Housekeeping. This repo is 39.4 GB, of which 36.35 GB is five intermediate step*.pt checkpoints that nothing loads. Only final.pt and the ONNX export are used.

Use it

No loading snippet for this model yet.

Training data

Trained on the Ghana Speech dataset and related Ghanaian corpora, licensed CC BY-NC 4.0.

Intended use & license

Non-commercial use only (CC BY-NC 4.0). This is inherited from the training data and required by the terms under which the compute was granted: models trained in that window are non-commercial by condition of access, not by inference.

Limitations, stated plainly

  • Dagbani did get a recogniser, and the claim that it could not was wrong twice over. Every card on this account used to say that "one fine-tuning session on 74 validation rows would not change that". Those 74 rows are the eng-dag machine-translation validation split; the Dagbani speech data in this same account is waxal_dag — 13,228 training rows, 1,750 validation rows, ~71 hours, 1,041 speakers with the largest at 1%. Trained on it, tekyerema-asr-mms-dag scores 36.94 / 11.71, against the 86.59 / 33.95 this project had believed was the ceiling. It still loses to FarmerlineML/w2v-bert-2.0_2026_dagbani_ASR at 29.20 / 9.27, which is what the agent actually serves. A number carried across from a translation table into a speech claim was then repeated on every card here until 2026-09-22.
  • Evaluation is on read and machine-translated text. No recordings of people speaking agent commands in these languages exist. Numbers measured this way are optimistic about phrasing and pessimistic about code-switching, and should not be read as field performance.
  • Research work from a hackathon entry, not a supported product.

The rest of the family

Recognisers

Voices

Agent models

Translation

Routing

Acknowledgements

Compute resources provided by AI Skills and Compute Africa (AISCA). Trained on the Ghana NLP H200 GPU. Please keep derivatives non-commercial and share improvements back with the Ghana NLP community (ghananlpcommunity).

Downloads last month
49
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using PrinceAlhassanNasamu/kasa42-asr 1