Fun-CosyVoice3-0.5B-2512 — ONNX

ONNX export of FunAudioLLM/Fun-CosyVoice3-0.5B-2512 (Alibaba Tongyi Lab / FunAudioLLM, Apache-2.0) — zero-shot multilingual voice cloning — split into graphs that run with plain ONNX Runtime and a small host-side loop. Exported for WinSTT's local read-aloud engine (a Rust port of the host loop); all credit for the model goes to the CosyVoice authors (paper, code).

The LLM is the RL-tuned checkpoint (llm.rl.pt; upstream reports en WER 1.68 vs 2.24 and zh CER 0.81 vs 1.21 for the SFT one).

Files

file from contents
tokenizer.json this repo Qwen2 BPE + CosyVoice3 special tokens (`<
text_embedding_fp16.onnx this repo Qwen2 input embedding table (fp16 storage, fp32 out) (272.3 MB)
speech_embedding.onnx this repo speech-token embedding (6761 × 896) (24.2 MB)
llm.onnx + .data this repo Qwen2 LM + speech head, prefill/decode with KV cache, fp32 (1,456.8 MB)
llm_int8.onnx + .data this repo same, dynamic int8 MatMul (per-channel) — recommended CPU default (367.3 MB)
llm_q4.onnx + .data this repo same, MatMulNBits 4-bit (block 32, symmetric) (228.7 MB)
flow_encoder.onnx this repo flow token embedding + pre-lookahead + 2× upsample, speaker projection (4.5 MB)
hift.onnx this repo HiFT vocoder (f0 predictor, harmonic source, generator) up to STFT magnitude/phase (83.4 MB)
voices/zh-female.wav this repo built-in voice: upstream asset/zero_shot_prompt.wav (Apache-2.0) (0.3 MB)
voices/en-male.wav this repo built-in voice: LibriTTS-R test-clean 8224_274384_000016_000000 (CC BY 4.0) (0.4 MB)

Upstream already publishes usable ONNX for the speaker encoder, speech tokenizer and DiT estimator, so those are not republished — fetch them from the upstream repo: campplus.onnx, speech_tokenizer_v3.onnx, flow.decoder.estimator.fp32.onnx.

Graph contract (batch = 1 unless noted)

graph inputs outputs
text_embedding_fp16 input_ids i64 [1,S] inputs_embeds f32 [1,S,896] (fp16 table, cast in-graph)
speech_embedding input_ids i64 [1,S] inputs_embeds f32 [1,S,896]
llm* inputs_embeds f32 [1,S,896], attention_mask i64 [1,past+S], position_ids i64 [1,S], past_key_values.{0..23}.{key,value} f32 [1,2,past,64] logits f32 [1,6761] (last position only), present.{0..23}.{key,value} [1,2,past+S,64]
flow_encoder token i64 [1,T] (prompt tokens ++ new tokens), embedding f32 [1,192] mu f32 [1,80,2T], spks f32 [1,80]
flow.decoder.estimator.fp32 (upstream) x,mu,cond f32 [2,80,F], mask [2,1,F], t [2], spks [2,80] velocity [2,80,F]
hift speech_feat f32 [1,80,F], noise f32 [1,480F,9] uniform [0,1) magnitude, phase f32 [1,9,120F+1]
campplus (upstream) Kaldi fbank [1,T,80] (mean-normalised) speaker embedding [1,192]
speech_tokenizer_v3 (upstream) Whisper 128-bin log-mel [1,128,T], length i32 [1] speech tokens [1,T/4]

Host side (what the graphs deliberately leave out):

  1. Prompt features from the reference clip (≤ 30 s): speech tokens (16 kHz Whisper log-mel → tokenizer), CAM++ embedding (16 kHz Kaldi fbank, 80 bins, dither 0, minus per-utterance mean), and the 24 kHz Matcha mel (n_fft 1920, hop 480, 80 Slaney mels, fmax 12 kHz, ln(clamp(1e-5))). Trim so mel_frames == 2 * tokens.
  2. LLM prompt: [speech_emb[6561] (sos), text_emb(prompt_text ++ text), speech_emb[6563] (task), speech_emb(prompt tokens)].
    • zero-shot: prompt_text = "You are a helpful assistant.<|endofprompt|>" + transcript, prompt tokens = reference tokens.
    • cross-lingual (no transcript): text = "You are a helpful assistant.<|endofprompt|>" + text, no prompt text/tokens.
    • instruct: prompt_text = "You are a helpful assistant. <instruction><|endofprompt|>", no prompt tokens.
    • WinSTT also uses cross-lingual whenever the transcript's script differs from the text's (Han / Kana / Hangul / Cyrillic / Latin): zero-shot with a Mandarin prompt reading German gave 25% WER, cross-lingual ~2–5%.
  3. Decode with Repetition-Aware Sampling (top-p 0.8, top-k 25, window 10, τ 0.1); stop on any id ≥ 6561; min/max length = 2×/20× the tts-text token count (EOS masked below min). Next input = speech_embedding[token]. At most 5 consecutive silent tokens (1,2,28,29,55,248,494,2241,2242,2322,2323) are passed on to the flow.
  4. Flow matching: x0 ~ N(0,1) [1,80,F]; cond = prompt mel then zeros; 10 Euler steps on t = 1 - cos(linspace(0,1,11)·π/2); batch 2 = (cond, zeros for mu/spks/cond); v = 1.7·v0 − 0.7·v1. Keep frames after the prompt mel.
  5. Vocoder: hift → istft(n_fft 16, hop 4, Hann periodic, center) of clip(mag,≤100)·e^{i·phase}, clamp ±0.99. Speed: linearly resample the mel to F/speed frames first.

noise replaces upstream's fixed SineGen2.sine_waves buffer (pass torch.rand / any U[0,1) noise); the f0 predictor runs in fp32 instead of upstream's fp64.

Parity (vs the PyTorch modules)

check max abs cosine
text_embedding_fp16 8.84e-05 1.000000
speech_embedding 0.00e+00 1.000000
llm.onnx prefill logp 1.25e-05 1.000000
llm.onnx decode logp (30 steps) 1.74e-05 1.000000
llm_int8.onnx prefill logp 3.01e+00 0.998772
llm_int8.onnx decode logp (30 steps) 2.31e+00 0.999591
llm_q4.onnx prefill logp 1.41e+00 0.999766
llm_q4.onnx decode logp (30 steps) 1.53e+00 0.999794
flow_encoder mu 1.07e-05 1.000000
flow_encoder spks 1.19e-07 1.000000
flow mel (flow.decoder.estimator.fp32.onnx) 1.58e-02 1.000000
hift + rust-style iSTFT (audio) 1.21e-02 0.999996

Teacher-forced decode (30 greedy steps):

  • llm.onnx top1_agree: 1
  • llm.onnx mean_KL: 3.65e-06
  • llm_int8.onnx top1_agree: 1
  • llm_int8.onnx mean_KL: 0.02164
  • llm_q4.onnx top1_agree: 1
  • llm_q4.onnx mean_KL: 0.006794

End-to-end quality (ONNX pipeline, this export)

LLM device lang WER / CER (zh) speaker sim RTF (whole run)
q4 Auto en 1.75% 0.924 1.21
q4 Auto zh 0.63% 0.910 1.21
q4 Auto de 2.27% 0.936 1.21
q4 Auto es 3.23% 0.927 1.21
int8 Auto en 1.75% 0.927 1.06
int8 Auto zh 0.63% 0.908 1.06
int8 Auto de 2.27% 0.945 1.06
int8 Auto es 3.23% 0.927 1.06

Rendered by WinSTT's Rust engine (release build) from these graphs: 10 sentences per language, alternating the two bundled voices, one sample each (RAS sampling, fixed seed). ASR = Whisper large-v3-turbo (zh scored as CER after zh-Hans conversion); similarity = WavLM-base-plus-sv cosine to the reference clip. RTF = wall time / audio seconds; Auto = DirectML on an RTX 3080 Ti for the DiT estimator + HiFT, with the LLM on 8 CPU threads (i9-12900KF). CPU-only (int8, all graphs on CPU) measured RTF ≈ 5 on an idle machine (Python reference loop) — the 10-step CFG DiT dominates.

License

Apache-2.0, same as the upstream model. Bundled reference voices: voices/zh-female.wav is upstream's Apache-2.0 prompt clip; voices/en-male.wav is from LibriTTS-R (CC BY 4.0, Koizumi et al. 2023; source LibriVox, public domain).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Masterx/Fun-CosyVoice3-0.5B-2512-ONNX

Quantized
(22)
this model

Paper for Masterx/Fun-CosyVoice3-0.5B-2512-ONNX