Instructions to use dys-asr/parakeet-tdt-0.6b-sapc1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dys-asr/parakeet-tdt-0.6b-sapc1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="dys-asr/parakeet-tdt-0.6b-sapc1")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("dys-asr/parakeet-tdt-0.6b-sapc1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
parakeet-tdt-0.6b-sapc1
nvidia/parakeet-tdt-0.6b-v3 fine-tuned on the Speech Accessibility Project
corpus 1 (SAPC1) train split. A token-and-duration transducer, not CTC:
it has a prediction network, a joint network and a duration head, and decodes
autoregressively.
On the full SAPC1 dev split, against the CTC sibling trained on the same corpus:
| WER | CER | |
|---|---|---|
dys-asr/parakeet-ctc-0.6b-sapc1 |
10.32% | 6.09% |
dys-asr/parakeet-tdt-0.6b-sapc1 |
9.38% | 5.84% |
Both are all 31,114 dev utterances, no language model, both sides through the same normaliser. The transducer is 0.94 WER points better.
It wins despite a handicap, not because of an advantage. Two parts of the CTC recipe could not be matched, and both cost the transducer rather than helping it:
- Its BatchNorm statistics span 16 examples against the CTC model's 32.
Parakeet normalises over the per-device batch in 24
BatchNorm1dlayers, and the transducer's joint network -- which materialises[batch, frames, labels, vocab]at an 8,192 vocabulary -- runs out of memory on an 80 GB A100 above per-device 8. Accumulation restores the gradient to an effective batch of 32 but cannot restore a batch statistic, which is computed per micro-batch. - It trains on 141 fewer utterances, because labels are capped at 130 tokens. One 341-token transcript inflates the joint tensor 3.3x and exhausts the card.
Requirements
transformers>=5.9. ParakeetForTDT does not exist before it, so this model
cannot be loaded by the 4.x releases that run the CTC siblings.
import torch
from transformers import AutoProcessor, ParakeetForTDT
model_id = "dys-asr/parakeet-tdt-0.6b-sapc1"
processor = AutoProcessor.from_pretrained(model_id)
model = ParakeetForTDT.from_pretrained(model_id).eval()
inputs = processor(audio, sampling_rate=16_000, return_tensors="pt") # 16 kHz mono
with torch.inference_mode():
outputs = model.generate(**inputs)
print(processor.batch_decode(outputs.sequences, skip_special_tokens=True))
Decoding is autoregressive rather than an argmax over frames, so it is substantially slower than the CTC siblings -- about 13x fewer samples per second per GPU during training, and a similar gap at inference.
Output text convention
This model writes numbers as words and emits lower-case, unpunctuated text.
audio: "lower the temperature three degrees"
output: lower the temperature three degrees # not "... 3 degrees"
Lower case is not cosmetic. v3's tokeniser holds 8,192 tokens and is
overwhelmingly lower-case -- only 740 contain any capital -- so upper-case labels
cost 4.06 tokens per word against 1.67, because words shatter into single
characters: WHAT becomes W, H, A, T. Training on that would teach the
model to emit characters instead of its native subwords, so labels are
lower-cased before tokenising. Scoring is unaffected, since both hypothesis and
reference are upper-cased by the same normaliser.
Numerals are verbalised as they are for the CTC siblings. If you score this model, apply the same convention to your references.
Base model
This is v3, not v2. nvidia/parakeet-tdt-0.6b-v2 ships only a .nemo
archive -- no config.json, no safetensors -- so from_pretrained cannot open
it and there is no NeMo-free way to fine-tune it. v3 publishes transformers
weights. It is the multilingual 25-language model rather than the
English-specialised v2, which is a real difference and is why this is not a
reproduction of any published v2 result.
Training data
SAPC1 train split, 0.5-30 s, labels capped at 130 tokens: 212,054 of 218,900 utterances.
| train | dev (evaluation) | |
|---|---|---|
| speakers | 580 | 83 |
| utterances used | 212,054 | 31,114 |
Training hyperparameters
Matched to parakeet-ctc-0.6b-sapc1 wherever a transducer allows it.
| base model | nvidia/parakeet-tdt-0.6b-v3 |
| optimizer | AdamW, weight decay 0.01 |
| learning rate | 1e-4 |
| schedule | tri-stage, 10% warmup, 40% hold |
| per-device batch | 8 x 2 GPUs, SyncBatchNorm over 16 |
| gradient accumulation | 2 |
| effective batch | 32 |
| epochs | 10 (66,320 optimizer updates) |
| precision | bf16 mixed, float32 weights |
| layerdrop | 0.05 |
| gradient clipping | 1.0 |
| max audio duration | 0.5-30 s |
| max label length | 130 tokens |
| gradient checkpointing | off |
| batch sampling | length-grouped |
| seed | 42 |
Weights stay float32 under bf16 autocast rather than being cast: loading in bfloat16 fails outright, because the feature extractor emits float32 and the first convolution rejects the mismatch.
Dev WER on a fixed 4,000-utterance subset, by epoch:
| epoch | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| WER | 11.47 | 10.76 | 10.56 | 10.41 | 10.27 | 10.03 | 9.99 | 9.50 | 9.42 | 9.42 |
Still falling at epoch 10, so the recipe has not saturated.
Layerdrop
An earlier reading of the training loss suggested layerdrop 0.05 was destructive here: about 7% of micro-batches score around 2000 against a median of 2.4, because dropping an encoder layer can leave the alignment lattice unsatisfiable in a way CTC tolerates and a transducer does not. Run head to head against layerdrop 0.0 for three epochs, the two were indistinguishable -- 11.47 against 11.52 at epoch 1, 10.51 against 10.48 at epoch 3 -- so the recipe's 0.05 was kept. A transducer's training loss is not comparable across regularisation settings, and reading it as if it were is what produced the wrong conclusion.
Limitations
- Requires
transformers>=5.9. Not loadable by the 4.x line. - Slow relative to CTC. Autoregressive decoding, roughly an order of magnitude fewer samples per second.
- Lower-case, unpunctuated, numerals as words.
- Multilingual base fine-tuned on English only, discarding capability it had.
- No per-speaker breakdown published, unlike the CTC siblings: the final evaluation did not retain per-utterance predictions.
- Single training seed. No variance estimate.
- BatchNorm statistics span 16, not 32, unlike the CTC sibling. See above.
- Downloads last month
- -
Model tree for dys-asr/parakeet-tdt-0.6b-sapc1
Base model
nvidia/parakeet-tdt-0.6b-v3Evaluation results
- WER on SAPC1 dev (Speech Accessibility Project)self-reported9.380
- CER on SAPC1 dev (Speech Accessibility Project)self-reported5.840