Nabra-82M v0.1 β ONNX
ONNX exports of oddadmix/Nabra-82M-v0.1, a Kokoro-82M model fine-tuned for Modern Standard Arabic, packaged for on-device CPU inference.
All 81.76M parameters are intact. No layers removed, no structural surgery.
| File | Size | Precision |
|---|---|---|
nabra_fp32.onnx |
325.6 MB | FP32 weights, FP32 compute |
The FP16 variant now lives in nabra-82m-fp16-onnx.
Inputs and outputs
input_ids : int64 [1, sequence] phoneme ids, BOS/EOS wrapped (id 0)
ref_s : float32[1, 256] style vector from the voice pack
speed : float32[1] 1.0 = normal
-> audio : float32[samples] 24 kHz mono
Sequence length is capped at 510 tokens: the voice pack is [510, 1, 256],
so longer inputs have no style row to condition on. Split longer text on
sentence boundaries.
Usage
import numpy as np, onnxruntime as ort
sess = ort.InferenceSession("nabra_fp32.onnx", providers=["CPUExecutionProvider"])
audio = sess.run(None, {
"input_ids": np.asarray([input_ids], dtype=np.int64), # phoneme ids
"ref_s": ref_s.astype(np.float32).reshape(1, 256),
"speed": np.asarray([1.0], dtype=np.float32),
})[0].squeeze()
Phonemes come from espeak-ng (ar voice, IPA output), mapped through the
model's vocab.json. The graph's stochastic nodes are seeded, so a given
session produces reproducible output.
Measured performance
POCO F3 (Snapdragon 870), ONNX Runtime 1.26.0 arm64, 4 threads, all cores. Median of three interleaved rounds with the variant order rotated each round β the device heats measurably during a run, so ordering matters more than it looks.
| Variant | Size | RTF | Peak RSS |
|---|---|---|---|
nabra_fp32.onnx |
325.6 MB | 0.628 | 466 MB |
nabra_fp16.onnx |
163.4 MB | 0.622 | 530 MB |
RTF is the ratio of synthesis time to audio duration; below 1.0 is faster than real time.
Two things worth knowing before you pick
FP16 halves the file but uses more RAM. 530 MB against FP32's 466 MB. The weights are stored FP16 and cast to FP32 to compute, so the runtime holds both the FP16 initializers and the FP32 copies. Choose FP16 to reduce download and storage; do not choose it to reduce memory.
FP16 is not faster. 0.622 against 0.628 is inside the run-to-run spread of a single variant. This is FP16 storage, not FP16 arithmetic β ARMv8.0-A has no native FP16 math, and forcing it there is slower, not quicker.
If peak memory is your constraint, FP32 is the better choice of these two.
Building the FP16 variant
onnxconverter_common.float16.convert_float_to_float16 does not produce a
loadable model from this graph β it leaves nodes with mismatched FP16/FP32
inputs, because it does not reliably retype tensor-sequence values or recurse
into Loop / If subgraph bodies, all of which this vocoder uses.
The variant here is built instead by converting each float initializer to FP16
and inserting a Cast back to FP32 before its consumer. That loads, and it
keeps compute in FP32 where the hardware is fast.
Why there is no INT8 or INT4 variant here
Both were built and measured. Neither is worth shipping for this architecture, and the reasons are specific rather than general β they are about what this model is, not about quantization being bad.
INT8 dynamic β smallest file, 5.3Γ slower
onnxruntime.quantization.quantize_dynamic with default settings is the
one-liner most people reach for first. On this model it produces the smallest
artifact and the worst latency:
| Variant | Size | RTF | Peak RSS |
|---|---|---|---|
nabra_fp32.onnx |
325.6 MB | 0.628 | 466 MB |
| INT8 dynamic (not published) | 114.7 MB | 3.331 | 285 MB |
Measured 3.322 / 3.331 / 3.350 across three rounds β the tightest numbers in the whole comparison, so this is not noise.
The cause is ConvInteger. ONNX Runtime's integer convolution kernels are a
large regression on ARM64, and this model is 73% Conv by inference time.
MatMul does get a genuine fast integer kernel, but MatMul is only 3.9% of
runtime here, so restricting quantization to it buys almost nothing.
Note the trap: it is the smallest file and it has the lowest peak RSS, so every selection heuristic except measuring actual latency picks it.
Weight-only INT8 β works, but buys no speed
Per-output-channel weight-only INT8 (no activation quantization, no calibration data) does pass quality checks and roughly halves the file. It does not make inference faster, and the reason is worth knowing before trying it:
ONNX Runtime constant-folds DequantizeLinear on an initializer back to FP32
at session build. The weights are FP32 in memory from then on and the same
FP32 kernels run. Quantizing weights changes what you store, not what the CPU
does. Expect a download-size win; do not expect a latency win.
INT4 β 11Γ the weight error of INT8
INT4 runs fine on ONNX Runtime 1.26 at opset 21 and packs two values per byte. The blocker is numerical. Measured on this model's own generator weights:
| Scheme | median rel. error | max | bits/weight |
|---|---|---|---|
| INT8 per-channel | 0.0088 | 0.026 | 8.00 |
| INT4 per-channel | 0.1595 | 0.470 | 4.00 |
| INT4 block=64 | 0.1106 | 0.139 | 4.25 |
| INT4 block=32 | 0.0985 | 0.112 | 4.50 |
Block-wise scaling behaves differently here than it does in LLM quantization, and the difference explains the result: it cuts the maximum error hard (0.470 β 0.112) while barely moving the median (0.160 β 0.099). Blocking wins when a few outliers drag the scale for everything sharing it. These convolution kernels are near-Gaussian β there are no outliers to isolate, so every block inherits roughly the same range as the whole channel.
Also worth noting: at block=32 the real cost is 4.5 bits per weight once each block carries its own scale, so the win over INT8 is 1.78Γ, not 2Γ.
Post-hoc INT4 would need quantization-aware training to be usable.
The general lesson
Compute in a neural vocoder is set by how much work happens at audio rate, not by parameter count. On this model the linguistic front-end (PL-BERT plus the prosody predictor) runs at frame rate and costs 6% of inference; the vocoder runs at 24 kHz and costs 82%. Quantization reduces stored bytes; it does not move that boundary. Meaningful latency gains here require a smaller vocoder, not a smaller number format.
Limitations
- Arabic (MSA) only, single voice. Diacritized input works best.
- 510 token cap per call.
- Peak memory scales with utterance length, not just weights β roughly 470 MB for 3 s of audio and about 1 GB for 17 s on FP32. Synthesize sentence by sentence on memory-constrained devices.
- Benchmarked on one device (Snapdragon 870). Low-end SoCs without out-of-order cores are substantially slower: on the same phone restricted to its four Cortex-A55 cores, RTF is ~2.8.
License
Apache-2.0, following the base model.
Quantized variants
| Repo | Precision | Size |
|---|---|---|
| nabra-82m-fp16-onnx | FP16 | 163 MB |
| nabra-82m-int8-onnx | INT8 | 83 MB |
| nabra-82m-int4qat-onnx | INT4-QAT | 57 MB |
The FP32 weights in this repo serve as the fine-tune reference; the quantized repos are drop-in ONNX replacements for on-device deployment.