Instructions to use Masterx/Fun-CosyVoice3-0.5B-2512-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- CosyVoice
How to use Masterx/Fun-CosyVoice3-0.5B-2512-ONNX with CosyVoice:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Fun-CosyVoice3-0.5B-2512 — ONNX
ONNX export of FunAudioLLM/Fun-CosyVoice3-0.5B-2512 (Alibaba Tongyi Lab / FunAudioLLM, Apache-2.0) — zero-shot multilingual voice cloning — split into graphs that run with plain ONNX Runtime and a small host-side loop. Exported for WinSTT's local read-aloud engine (a Rust port of the host loop); all credit for the model goes to the CosyVoice authors (paper, code).
The LLM is the RL-tuned checkpoint (llm.rl.pt; upstream reports en WER 1.68 vs 2.24 and zh CER
0.81 vs 1.21 for the SFT one).
Files
| file | from | contents |
|---|---|---|
tokenizer.json |
this repo | Qwen2 BPE + CosyVoice3 special tokens (`< |
text_embedding_fp16.onnx |
this repo | Qwen2 input embedding table (fp16 storage, fp32 out) (272.3 MB) |
speech_embedding.onnx |
this repo | speech-token embedding (6761 × 896) (24.2 MB) |
llm.onnx + .data |
this repo | Qwen2 LM + speech head, prefill/decode with KV cache, fp32 (1,456.8 MB) |
llm_int8.onnx + .data |
this repo | same, dynamic int8 MatMul (per-channel) — recommended CPU default (367.3 MB) |
llm_q4.onnx + .data |
this repo | same, MatMulNBits 4-bit (block 32, symmetric) (228.7 MB) |
flow_encoder.onnx |
this repo | flow token embedding + pre-lookahead + 2× upsample, speaker projection (4.5 MB) |
hift.onnx |
this repo | HiFT vocoder (f0 predictor, harmonic source, generator) up to STFT magnitude/phase (83.4 MB) |
voices/zh-female.wav |
this repo | built-in voice: upstream asset/zero_shot_prompt.wav (Apache-2.0) (0.3 MB) |
voices/en-male.wav |
this repo | built-in voice: LibriTTS-R test-clean 8224_274384_000016_000000 (CC BY 4.0) (0.4 MB) |
Upstream already publishes usable ONNX for the speaker encoder, speech tokenizer and DiT estimator, so
those are not republished — fetch them from the upstream repo:
campplus.onnx, speech_tokenizer_v3.onnx, flow.decoder.estimator.fp32.onnx.
Graph contract (batch = 1 unless noted)
| graph | inputs | outputs |
|---|---|---|
text_embedding_fp16 |
input_ids i64 [1,S] |
inputs_embeds f32 [1,S,896] (fp16 table, cast in-graph) |
speech_embedding |
input_ids i64 [1,S] |
inputs_embeds f32 [1,S,896] |
llm* |
inputs_embeds f32 [1,S,896], attention_mask i64 [1,past+S], position_ids i64 [1,S], past_key_values.{0..23}.{key,value} f32 [1,2,past,64] |
logits f32 [1,6761] (last position only), present.{0..23}.{key,value} [1,2,past+S,64] |
flow_encoder |
token i64 [1,T] (prompt tokens ++ new tokens), embedding f32 [1,192] |
mu f32 [1,80,2T], spks f32 [1,80] |
flow.decoder.estimator.fp32 (upstream) |
x,mu,cond f32 [2,80,F], mask [2,1,F], t [2], spks [2,80] |
velocity [2,80,F] |
hift |
speech_feat f32 [1,80,F], noise f32 [1,480F,9] uniform [0,1) |
magnitude, phase f32 [1,9,120F+1] |
campplus (upstream) |
Kaldi fbank [1,T,80] (mean-normalised) | speaker embedding [1,192] |
speech_tokenizer_v3 (upstream) |
Whisper 128-bin log-mel [1,128,T], length i32 [1] | speech tokens [1,T/4] |
Host side (what the graphs deliberately leave out):
- Prompt features from the reference clip (≤ 30 s): speech tokens (16 kHz Whisper log-mel → tokenizer),
CAM++ embedding (16 kHz Kaldi fbank, 80 bins, dither 0, minus per-utterance mean), and the 24 kHz
Matcha mel (n_fft 1920, hop 480, 80 Slaney mels, fmax 12 kHz,
ln(clamp(1e-5))). Trim somel_frames == 2 * tokens. - LLM prompt:
[speech_emb[6561] (sos), text_emb(prompt_text ++ text), speech_emb[6563] (task), speech_emb(prompt tokens)].- zero-shot:
prompt_text = "You are a helpful assistant.<|endofprompt|>" + transcript, prompt tokens = reference tokens. - cross-lingual (no transcript):
text = "You are a helpful assistant.<|endofprompt|>" + text, no prompt text/tokens. - instruct:
prompt_text = "You are a helpful assistant. <instruction><|endofprompt|>", no prompt tokens. - WinSTT also uses cross-lingual whenever the transcript's script differs from the text's (Han / Kana / Hangul / Cyrillic / Latin): zero-shot with a Mandarin prompt reading German gave 25% WER, cross-lingual ~2–5%.
- zero-shot:
- Decode with Repetition-Aware Sampling (top-p 0.8, top-k 25, window 10, τ 0.1); stop on any id ≥ 6561;
min/max length = 2×/20× the tts-text token count (EOS masked below min). Next input =
speech_embedding[token]. At most 5 consecutive silent tokens (1,2,28,29,55,248,494,2241,2242,2322,2323) are passed on to the flow. - Flow matching:
x0 ~ N(0,1)[1,80,F];cond= prompt mel then zeros; 10 Euler steps ont = 1 - cos(linspace(0,1,11)·π/2); batch 2 = (cond, zeros for mu/spks/cond);v = 1.7·v0 − 0.7·v1. Keep frames after the prompt mel. - Vocoder:
hift→istft(n_fft 16, hop 4, Hann periodic, center)ofclip(mag,≤100)·e^{i·phase}, clamp ±0.99. Speed: linearly resample the mel toF/speedframes first.
noise replaces upstream's fixed SineGen2.sine_waves buffer (pass torch.rand / any U[0,1) noise); the
f0 predictor runs in fp32 instead of upstream's fp64.
Parity (vs the PyTorch modules)
| check | max abs | cosine |
|---|---|---|
| text_embedding_fp16 | 8.84e-05 | 1.000000 |
| speech_embedding | 0.00e+00 | 1.000000 |
| llm.onnx prefill logp | 1.25e-05 | 1.000000 |
| llm.onnx decode logp (30 steps) | 1.74e-05 | 1.000000 |
| llm_int8.onnx prefill logp | 3.01e+00 | 0.998772 |
| llm_int8.onnx decode logp (30 steps) | 2.31e+00 | 0.999591 |
| llm_q4.onnx prefill logp | 1.41e+00 | 0.999766 |
| llm_q4.onnx decode logp (30 steps) | 1.53e+00 | 0.999794 |
| flow_encoder mu | 1.07e-05 | 1.000000 |
| flow_encoder spks | 1.19e-07 | 1.000000 |
| flow mel (flow.decoder.estimator.fp32.onnx) | 1.58e-02 | 1.000000 |
| hift + rust-style iSTFT (audio) | 1.21e-02 | 0.999996 |
Teacher-forced decode (30 greedy steps):
- llm.onnx top1_agree: 1
- llm.onnx mean_KL: 3.65e-06
- llm_int8.onnx top1_agree: 1
- llm_int8.onnx mean_KL: 0.02164
- llm_q4.onnx top1_agree: 1
- llm_q4.onnx mean_KL: 0.006794
End-to-end quality (ONNX pipeline, this export)
| LLM | device | lang | WER / CER (zh) | speaker sim | RTF (whole run) |
|---|---|---|---|---|---|
| q4 | Auto | en | 1.75% | 0.924 | 1.21 |
| q4 | Auto | zh | 0.63% | 0.910 | 1.21 |
| q4 | Auto | de | 2.27% | 0.936 | 1.21 |
| q4 | Auto | es | 3.23% | 0.927 | 1.21 |
| int8 | Auto | en | 1.75% | 0.927 | 1.06 |
| int8 | Auto | zh | 0.63% | 0.908 | 1.06 |
| int8 | Auto | de | 2.27% | 0.945 | 1.06 |
| int8 | Auto | es | 3.23% | 0.927 | 1.06 |
Rendered by WinSTT's Rust engine (release build) from these graphs: 10 sentences per language, alternating the two bundled voices, one sample each (RAS sampling, fixed seed). ASR = Whisper large-v3-turbo (zh scored as CER after zh-Hans conversion); similarity = WavLM-base-plus-sv cosine to the reference clip. RTF = wall time / audio seconds; Auto = DirectML on an RTX 3080 Ti for the DiT estimator + HiFT, with the LLM on 8 CPU threads (i9-12900KF). CPU-only (int8, all graphs on CPU) measured RTF ≈ 5 on an idle machine (Python reference loop) — the 10-step CFG DiT dominates.
License
Apache-2.0, same as the upstream model. Bundled reference voices: voices/zh-female.wav is upstream's Apache-2.0 prompt clip; voices/en-male.wav is from LibriTTS-R (CC BY 4.0, Koizumi et al. 2023; source LibriVox, public domain).
Model tree for Masterx/Fun-CosyVoice3-0.5B-2512-ONNX
Base model
FunAudioLLM/Fun-CosyVoice3-0.5B-2512