You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

parakeet-tdt-0.6b-sapc1

nvidia/parakeet-tdt-0.6b-v3 fine-tuned on the Speech Accessibility Project corpus 1 (SAPC1) train split. A token-and-duration transducer, not CTC: it has a prediction network, a joint network and a duration head, and decodes autoregressively.

On the full SAPC1 dev split, against the CTC sibling trained on the same corpus:

WER CER
dys-asr/parakeet-ctc-0.6b-sapc1 10.32% 6.09%
dys-asr/parakeet-tdt-0.6b-sapc1 9.38% 5.84%

Both are all 31,114 dev utterances, no language model, both sides through the same normaliser. The transducer is 0.94 WER points better.

It wins despite a handicap, not because of an advantage. Two parts of the CTC recipe could not be matched, and both cost the transducer rather than helping it:

  • Its BatchNorm statistics span 16 examples against the CTC model's 32. Parakeet normalises over the per-device batch in 24 BatchNorm1d layers, and the transducer's joint network -- which materialises [batch, frames, labels, vocab] at an 8,192 vocabulary -- runs out of memory on an 80 GB A100 above per-device 8. Accumulation restores the gradient to an effective batch of 32 but cannot restore a batch statistic, which is computed per micro-batch.
  • It trains on 141 fewer utterances, because labels are capped at 130 tokens. One 341-token transcript inflates the joint tensor 3.3x and exhausts the card.

Requirements

transformers>=5.9. ParakeetForTDT does not exist before it, so this model cannot be loaded by the 4.x releases that run the CTC siblings.

import torch
from transformers import AutoProcessor, ParakeetForTDT

model_id = "dys-asr/parakeet-tdt-0.6b-sapc1"
processor = AutoProcessor.from_pretrained(model_id)
model = ParakeetForTDT.from_pretrained(model_id).eval()

inputs = processor(audio, sampling_rate=16_000, return_tensors="pt")  # 16 kHz mono
with torch.inference_mode():
    outputs = model.generate(**inputs)
print(processor.batch_decode(outputs.sequences, skip_special_tokens=True))

Decoding is autoregressive rather than an argmax over frames, so it is substantially slower than the CTC siblings -- about 13x fewer samples per second per GPU during training, and a similar gap at inference.

Output text convention

This model writes numbers as words and emits lower-case, unpunctuated text.

audio:  "lower the temperature three degrees"
output: lower the temperature three degrees     # not "... 3 degrees"

Lower case is not cosmetic. v3's tokeniser holds 8,192 tokens and is overwhelmingly lower-case -- only 740 contain any capital -- so upper-case labels cost 4.06 tokens per word against 1.67, because words shatter into single characters: WHAT becomes W, H, A, T. Training on that would teach the model to emit characters instead of its native subwords, so labels are lower-cased before tokenising. Scoring is unaffected, since both hypothesis and reference are upper-cased by the same normaliser.

Numerals are verbalised as they are for the CTC siblings. If you score this model, apply the same convention to your references.

Base model

This is v3, not v2. nvidia/parakeet-tdt-0.6b-v2 ships only a .nemo archive -- no config.json, no safetensors -- so from_pretrained cannot open it and there is no NeMo-free way to fine-tune it. v3 publishes transformers weights. It is the multilingual 25-language model rather than the English-specialised v2, which is a real difference and is why this is not a reproduction of any published v2 result.

Training data

SAPC1 train split, 0.5-30 s, labels capped at 130 tokens: 212,054 of 218,900 utterances.

train dev (evaluation)
speakers 580 83
utterances used 212,054 31,114

Training hyperparameters

Matched to parakeet-ctc-0.6b-sapc1 wherever a transducer allows it.

base model nvidia/parakeet-tdt-0.6b-v3
optimizer AdamW, weight decay 0.01
learning rate 1e-4
schedule tri-stage, 10% warmup, 40% hold
per-device batch 8 x 2 GPUs, SyncBatchNorm over 16
gradient accumulation 2
effective batch 32
epochs 10 (66,320 optimizer updates)
precision bf16 mixed, float32 weights
layerdrop 0.05
gradient clipping 1.0
max audio duration 0.5-30 s
max label length 130 tokens
gradient checkpointing off
batch sampling length-grouped
seed 42

Weights stay float32 under bf16 autocast rather than being cast: loading in bfloat16 fails outright, because the feature extractor emits float32 and the first convolution rejects the mismatch.

Dev WER on a fixed 4,000-utterance subset, by epoch:

epoch 1 2 3 4 5 6 7 8 9 10
WER 11.47 10.76 10.56 10.41 10.27 10.03 9.99 9.50 9.42 9.42

Still falling at epoch 10, so the recipe has not saturated.

Layerdrop

An earlier reading of the training loss suggested layerdrop 0.05 was destructive here: about 7% of micro-batches score around 2000 against a median of 2.4, because dropping an encoder layer can leave the alignment lattice unsatisfiable in a way CTC tolerates and a transducer does not. Run head to head against layerdrop 0.0 for three epochs, the two were indistinguishable -- 11.47 against 11.52 at epoch 1, 10.51 against 10.48 at epoch 3 -- so the recipe's 0.05 was kept. A transducer's training loss is not comparable across regularisation settings, and reading it as if it were is what produced the wrong conclusion.

Limitations

  • Requires transformers>=5.9. Not loadable by the 4.x line.
  • Slow relative to CTC. Autoregressive decoding, roughly an order of magnitude fewer samples per second.
  • Lower-case, unpunctuated, numerals as words.
  • Multilingual base fine-tuned on English only, discarding capability it had.
  • No per-speaker breakdown published, unlike the CTC siblings: the final evaluation did not retain per-utterance predictions.
  • Single training seed. No variance estimate.
  • BatchNorm statistics span 16, not 32, unlike the CTC sibling. See above.
Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dys-asr/parakeet-tdt-0.6b-sapc1

Finetuned
(92)
this model

Evaluation results

  • WER on SAPC1 dev (Speech Accessibility Project)
    self-reported
    9.380
  • CER on SAPC1 dev (Speech Accessibility Project)
    self-reported
    5.840