Instructions to use Arm/pocket-tts-pt-mix-precision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use Arm/pocket-tts-pt-mix-precision with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("Arm/pocket-tts-pt-mix-precision") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - LiteRT
How to use Arm/pocket-tts-pt-mix-precision with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Arm/pocket-tts-pt-mix-precision with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Arm/pocket-tts-pt-mix-precision # Run inference directly in the terminal: llama cli -hf Arm/pocket-tts-pt-mix-precision
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Arm/pocket-tts-pt-mix-precision # Run inference directly in the terminal: llama cli -hf Arm/pocket-tts-pt-mix-precision
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Arm/pocket-tts-pt-mix-precision # Run inference directly in the terminal: ./llama-cli -hf Arm/pocket-tts-pt-mix-precision
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Arm/pocket-tts-pt-mix-precision # Run inference directly in the terminal: ./build/bin/llama-cli -hf Arm/pocket-tts-pt-mix-precision
Use Docker
docker model run hf.co/Arm/pocket-tts-pt-mix-precision
- LM Studio
- Jan
- Ollama
How to use Arm/pocket-tts-pt-mix-precision with Ollama:
ollama run hf.co/Arm/pocket-tts-pt-mix-precision
- Unsloth Desktop
- Docker Model Runner
How to use Arm/pocket-tts-pt-mix-precision with Docker Model Runner:
docker model run hf.co/Arm/pocket-tts-pt-mix-precision
- Lemonade
How to use Arm/pocket-tts-pt-mix-precision with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Arm/pocket-tts-pt-mix-precision
Run and chat with the model
lemonade run user.pocket-tts-pt-mix-precision-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
Pocket TTS Portuguese - Mixed Precision and Mixed Frameworks
This repository packages a mixed-framework deployment bundle derived from
kyutai/pocket-tts for Portuguese
text-to-speech. The bundle uses SentencePiece, GGUF, ONNX, and LiteRT artifacts
and supports streaming synthesis at 24 kHz.
โจ Key Highlights
- Faster than real-time reference โ a standard-size package was measured using one Armยฎ CPU
core with SME2:
- RTF is 0.222 on an Androidโข Vivo X300 smartphone.
- Low-latency reference โ measured with the standard-size package on an Androidโข Vivo X300
smartphone:
- Median end-to-end latency is 0.897 s for a prompt that produces 4.04 seconds of audio.
- Median time to first audio chunk is 0.107 s.
- Memory-efficient reference โ measured with the standard-size package:
- Peak process RSS is 820.3 MiB on an Androidโข Vivo X300 smartphone.
- Compact deployment โ the packaged model files total 382.9 MB.
- Portuguese text-to-speech โ generates spoken Portuguese from a text prompt.
- Packaged voice profiles โ includes 26 voice profiles.
- Mixed framework deployment โ combines SentencePiece, llama.cpp, ONNX Runtime, and LiteRT.
- Selective FP16 weights โ the audio decoder uses FP16 fully connected weights.
- Armยฎ CPU deployment โ optimized for efficient execution on Armยฎ CPUs.
- Streaming output โ returns synthesized speech at 24 kHz in streaming mode.
๐ฆ Model Details
Model Description
Pocket TTS is a lightweight text-to-speech model developed by Kyutai for efficient CPU execution.
- Developed by: Kyutai
- Model type: Portuguese text-to-speech with packaged voice profiles
- License: CC BY 4.0
- Base model:
kyutai/pocket-tts - Package form: SentencePiece, GGUF, ONNX, TFLite, and NumPy artifacts
Model Sources
- Base model: https://huggingface.co/kyutai/pocket-tts
- Upstream repository: https://github.com/kyutai-labs/pocket-tts
- Voice collection: https://huggingface.co/kyutai/tts-voices
๐ Get Started with the Model
๐ Compute Flow โ Early Access
The inference engine for this model package is available through the Compute Flow Early Access Program.
Want to try it?
๐ฉ Contact us at ai-early-access@arm.com to request access.
๐ Quality evaluation
Accuracy results for this mixed-precision package will be provided in a future update. However, no intelligibility regression was observed against the upstream FP32 implementation.
๐ฏ Performance evaluation
The benchmark figures below were obtained with the standard-size English package. The Portuguese variant uses the same runtime architecture and tensor dimensions, so similar performance is expected when prompts produce comparable output durations.
Performance was measured under the following conditions:
- One Armยฎ CPU thread.
- 5 warmups followed by 30 consecutive measured runs with no pause between runs.
- The Androidโข Vivo X300 smartphone screen was kept on.
- The model remained loaded and its state was reset between runs.
The following methodology and definitions were used:
- Runtime: LiteRT, llama.cpp, ONNX Runtime, and SentencePiece on CPU, with XNNPACK and KleidiAI.
- Input prompt: "Hello everyone. I am Jack and I am your personal assistant."
- Voice profile:
alba, the English package default. - Generated output: 4.04 seconds of 24 kHz audio per run.
- End-to-end latency is the summed model execution time for the complete utterance and excludes model setup.
- Time to first audio chunk is the elapsed execution time until the first streaming audio chunk is returned.
- Average memory is the mean process RSS sampled throughout setup and inference.
- Peak memory is the maximum sampled process high-water mark.
- Latency and memory were collected in separate executions to prevent memory sampling from affecting latency.
| Metric | English reference: Androidโข Vivo X300 |
|---|---|
| Model size | 382.9 MB |
| RTF | 0.222 |
| End-to-end latency, p50 | 0.897 s |
| End-to-end latency, p90 | 0.904 s |
| End-to-end latency, p99 | 0.906 s |
| Time to first audio chunk | 0.107 s |
| Peak memory | 820.3 MiB |
| Average memory | 811.0 MiB |
RTF is the total inference time divided by the generated audio duration, so lower values indicate faster processing and values below 1 indicate faster-than-real-time generation.
๐ ๏ธ Technical Specifications
Objective
Generate streaming 24 kHz Portuguese speech from a text prompt.
Runtime Architecture
| Component role | Framework / format |
|---|---|
| Text tokenization | SentencePiece |
| Text encoding, projection, and denoising | ONNX Runtime / ONNX |
| Flow language modeling | llama.cpp / GGUF |
| Audio decoding | LiteRT / TFLite |
| Voice conditioning | NumPy profile |
Precision and Quantization
The LiteRT audio decoder uses FP16 fully connected weights.
Input Specification
| Input | Description |
|---|---|
| Text prompt | UTF-8 string |
Output Specification
The model returns synthesized 24 kHz audio in streaming mode.
Repository Contents
pocket_tts_manifest.jsonโ model package manifest.tokenizer.modelโ SentencePiece tokenizer.*.onnx,*.gguf, and*.tfliteโ model components.assets/*.npyโ packaged voice profiles and model assets.metadata.yamlโ model metadata.SHA256SUMSโ model-bundle checksums for reproducibility.
๐๏ธ Model and Asset Origin
๐ Checksums
SHA256SUMS was generated by recursively hashing every regular file in the
model bundle, including files in subdirectories, except the generated root
SHA256SUMS and paths with a dotfile component.
From the model bundle root, verify the checked-out files with:
shasum -a 256 -c SHA256SUMS
- Downloads last month
- 10
We're not able to determine the quantization variants.
Model tree for Arm/pocket-tts-pt-mix-precision
Base model
kyutai/pocket-tts