Instructions to use Arm/whisper-large-v3-quantized.w4a8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Arm/whisper-large-v3-quantized.w4a8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Arm/whisper-large-v3-quantized.w4a8")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("Arm/whisper-large-v3-quantized.w4a8") model = AutoModelForSpeechSeq2Seq.from_pretrained("Arm/whisper-large-v3-quantized.w4a8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Whisper Large v3 W4A8
Model Overview
- Model architecture: Whisper large-v3
- Input: Audio
- Output: Text
- Model optimizations:
- Weight quantization: INT4
- Activation quantization: INT8
- Base model:
openai/whisper-large-v3
This model is a W4A8 GPTQ-quantized version of
openai/whisper-large-v3.
It is intended for automatic speech recognition evaluation and inference on
Arm CPU systems using vLLM.
The quantization workflow uses llmcompressor
to quantize Whisper linear layers with 4-bit integer weights and 8-bit integer
input activations. The language-model projection head is left unquantized.
The model was quantized without additional fine-tuning.
Intended Use
This model is intended for:
- Automatic speech recognition with Whisper-compatible tooling.
- Experiments comparing BF16 and compressed Whisper inference.
- Arm CPU vLLM deployments where reduced model size is useful.
This model is not intended to improve transcription quality over the base model. Users should validate quality, latency, memory use, and supported runtime behavior on their target hardware and workload before deployment.
Deployment
Use with vLLM
from vllm import LLM, SamplingParams
from vllm.assets.audio import AudioAsset
llm = LLM(
model="Arm/whisper-large-v3-quantized.w4a8",
max_model_len=448,
max_num_seqs=400,
limit_mm_per_prompt={"audio": 1},
)
inputs = {
"encoder_prompt": {
"prompt": "",
"multi_modal_data": {
"audio": AudioAsset("winning_call").audio_and_sample_rate,
},
},
"decoder_prompt": "<|startoftranscript|>",
}
outputs = llm.generate(inputs, SamplingParams(temperature=0.0, max_tokens=128))
print(outputs[0].outputs[0].text)
Quantization Details
| Field | Value |
|---|---|
| Base model | openai/whisper-large-v3 |
| Quantization method | GPTQ |
| Weight precision | INT4 |
| Weight strategy | Per-channel, symmetric |
| Input activation precision | INT8 |
| Activation strategy | Dynamic, per-token, symmetric |
| Quantized modules | Linear layers |
| Unquantized modules | proj_out |
| Calibration dataset | MLCommons/peoples_speech, subset test, split test |
| Calibration task prefix | English transcription |
| Calibration samples used for reported run | 1,024 |
| Maximum calibration sequence length | 2,048 |
Evaluation
Evaluation was run with lmms-eval
using the whisper_vllm model interface on LibriSpeech and FLEURS. Lower WER is
better.
| Benchmark | Split | BF16 WER | W4A8 WER | BF16/W4A8 Recovery |
|---|---|---|---|---|
| LibiriSpeech (WER) | test-clean | 2.1517 | 2.1989 | 97.9% |
| LibiriSpeech (WER) | test-other | 3.9352 | 4.0865 | 96.3% |
| Fleurs (WER) | cmn_hans_cn | 7.7907 | 8.1741 | 95.3% |
| Fleurs (WER) | en | 4.0442 | 4.0785 | 99.2% |
On average our INT4 implementation is able to recover 97.2% of BF16 WER.
Reproduce Quantization
Create a fresh quantization environment:
python -m venv .quantize
source .quantize/bin/activate
pip install llmcompressor
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu
Run quantization:
python quantize.py \
--model_path openai/whisper-large-v3 \
--save_dir data/model_dir \
--num_calibration_samples 1024
The output is written to:
data/model_dir/whisper-large-v3-quantized.w4a8
Reproduce Evaluation
Create a fresh evaluation environment:
python -m venv .eval
source .eval/bin/activate
export VLLM_VERSION=0.23.0
pip install "https://github.com/vllm-project/vllm/releases/download/v${VLLM_VERSION}/vllm-${VLLM_VERSION}+cpu-cp38-abi3-manylinux_2_34_aarch64.whl" --extra-index-url https://download.pytorch.org/whl/cpu
pip install editdistance
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu
git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git ~/lmms-eval
pip install -e ~/lmms-eval
Run the W4A8 evaluation:
lmms-eval \
--model whisper_vllm \
--model_args "pretrained=Arm/whisper-large-v3-quantized.w4a8" \
--tasks librispeech_test_other,librispeech_test_clean,fleurs \
--batch_size 64 \
--output_path results/w4a8/full_suite
Limitations
- Reported evaluations cover LibriSpeech and two FLEURS language splits only.
- The calibration samples use English transcription examples.
- Runtime support depends on vLLM, compressed-tensors, and target hardware.
- Quantization can change outputs, especially on languages, accents, domains, audio conditions, and decoding settings not covered by the reported evaluation.
Ethical Considerations
This model inherits the capabilities and limitations of Whisper large-v3. ASR systems can produce incorrect transcripts and may perform unevenly across languages, accents, dialects, speakers, domains, and recording conditions. Do not use transcripts as the sole basis for high-stakes decisions without human review.
About this version
This repository contains a W4A8 quantized version of OpenAI’s Whisper large-v3 model. Arm quantized the model using INT4 weight quantization and INT8 activation quantization to enable more efficient execution with vLLM on Arm-based platforms. No additional training or fine-tuning was applied by Arm. The original model architecture, intended automatic speech recognition and speech translation use cases, and known limitations remain applicable, although quantization may affect numerical behavior and accuracy.
Original model and documentation
For full details of the original model, please refer to the original OpenAI Whisper large-v3 model card: https://huggingface.co/openai/whisper-large-v3
Purpose of this release
Arm provides this quantized model to enable developers to evaluate and build applications using W4A8 Whisper large-v3 inference with vLLM on Arm-based systems. Users should validate its accuracy and behavior under the languages, audio conditions, decoding settings, and deployment environment relevant to their application.
- Downloads last month
- 48
Model tree for Arm/whisper-large-v3-quantized.w4a8
Base model
openai/whisper-large-v3