Pocket TTS Portuguese - Mixed Precision and Mixed Frameworks

This repository packages a mixed-framework deployment bundle derived from kyutai/pocket-tts for Portuguese text-to-speech. The bundle uses SentencePiece, GGUF, ONNX, and LiteRT artifacts and supports streaming synthesis at 24 kHz.

โœจ Key Highlights

  • Faster than real-time reference โ€” a standard-size package was measured using one Armยฎ CPU core with SME2:
    • RTF is 0.222 on an Androidโ„ข Vivo X300 smartphone.
  • Low-latency reference โ€” measured with the standard-size package on an Androidโ„ข Vivo X300 smartphone:
    • Median end-to-end latency is 0.897 s for a prompt that produces 4.04 seconds of audio.
    • Median time to first audio chunk is 0.107 s.
  • Memory-efficient reference โ€” measured with the standard-size package:
    • Peak process RSS is 820.3 MiB on an Androidโ„ข Vivo X300 smartphone.
  • Compact deployment โ€” the packaged model files total 382.9 MB.
  • Portuguese text-to-speech โ€” generates spoken Portuguese from a text prompt.
  • Packaged voice profiles โ€” includes 26 voice profiles.
  • Mixed framework deployment โ€” combines SentencePiece, llama.cpp, ONNX Runtime, and LiteRT.
  • Selective FP16 weights โ€” the audio decoder uses FP16 fully connected weights.
  • Armยฎ CPU deployment โ€” optimized for efficient execution on Armยฎ CPUs.
  • Streaming output โ€” returns synthesized speech at 24 kHz in streaming mode.

๐Ÿ“ฆ Model Details

Model Description

Pocket TTS is a lightweight text-to-speech model developed by Kyutai for efficient CPU execution.

  • Developed by: Kyutai
  • Model type: Portuguese text-to-speech with packaged voice profiles
  • License: CC BY 4.0
  • Base model: kyutai/pocket-tts
  • Package form: SentencePiece, GGUF, ONNX, TFLite, and NumPy artifacts

Model Sources

๐Ÿš€ Get Started with the Model

๐Ÿ”“ Compute Flow โ€” Early Access

The inference engine for this model package is available through the Compute Flow Early Access Program.

Want to try it?

๐Ÿ“ฉ Contact us at ai-early-access@arm.com to request access.

๐Ÿ“Š Quality evaluation

Accuracy results for this mixed-precision package will be provided in a future update. However, no intelligibility regression was observed against the upstream FP32 implementation.

๐ŸŽฏ Performance evaluation

The benchmark figures below were obtained with the standard-size English package. The Portuguese variant uses the same runtime architecture and tensor dimensions, so similar performance is expected when prompts produce comparable output durations.

Performance was measured under the following conditions:

  • One Armยฎ CPU thread.
  • 5 warmups followed by 30 consecutive measured runs with no pause between runs.
  • The Androidโ„ข Vivo X300 smartphone screen was kept on.
  • The model remained loaded and its state was reset between runs.

The following methodology and definitions were used:

  • Runtime: LiteRT, llama.cpp, ONNX Runtime, and SentencePiece on CPU, with XNNPACK and KleidiAI.
  • Input prompt: "Hello everyone. I am Jack and I am your personal assistant."
  • Voice profile: alba, the English package default.
  • Generated output: 4.04 seconds of 24 kHz audio per run.
  • End-to-end latency is the summed model execution time for the complete utterance and excludes model setup.
  • Time to first audio chunk is the elapsed execution time until the first streaming audio chunk is returned.
  • Average memory is the mean process RSS sampled throughout setup and inference.
  • Peak memory is the maximum sampled process high-water mark.
  • Latency and memory were collected in separate executions to prevent memory sampling from affecting latency.
Metric English reference: Androidโ„ข Vivo X300
Model size 382.9 MB
RTF 0.222
End-to-end latency, p50 0.897 s
End-to-end latency, p90 0.904 s
End-to-end latency, p99 0.906 s
Time to first audio chunk 0.107 s
Peak memory 820.3 MiB
Average memory 811.0 MiB

RTF is the total inference time divided by the generated audio duration, so lower values indicate faster processing and values below 1 indicate faster-than-real-time generation.

๐Ÿ› ๏ธ Technical Specifications

Objective

Generate streaming 24 kHz Portuguese speech from a text prompt.

Runtime Architecture

Component role Framework / format
Text tokenization SentencePiece
Text encoding, projection, and denoising ONNX Runtime / ONNX
Flow language modeling llama.cpp / GGUF
Audio decoding LiteRT / TFLite
Voice conditioning NumPy profile

Precision and Quantization

The LiteRT audio decoder uses FP16 fully connected weights.

Input Specification

Input Description
Text prompt UTF-8 string

Output Specification

The model returns synthesized 24 kHz audio in streaming mode.

Repository Contents

  • pocket_tts_manifest.json โ€” model package manifest.
  • tokenizer.model โ€” SentencePiece tokenizer.
  • *.onnx, *.gguf, and *.tflite โ€” model components.
  • assets/*.npy โ€” packaged voice profiles and model assets.
  • metadata.yaml โ€” model metadata.
  • SHA256SUMS โ€” model-bundle checksums for reproducibility.

๐Ÿ—‚๏ธ Model and Asset Origin

๐Ÿ” Checksums

SHA256SUMS was generated by recursively hashing every regular file in the model bundle, including files in subdirectories, except the generated root SHA256SUMS and paths with a dotfile component.

From the model bundle root, verify the checked-out files with:

shasum -a 256 -c SHA256SUMS
Downloads last month
10
GGUF
Model size
75.6M params
Architecture
gptneox
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Arm/pocket-tts-pt-mix-precision

Quantized
(53)
this model

Collection including Arm/pocket-tts-pt-mix-precision