Whisper Large v3 W4A8

Model Overview

  • Model architecture: Whisper large-v3
  • Input: Audio
  • Output: Text
  • Model optimizations:
    • Weight quantization: INT4
    • Activation quantization: INT8
  • Base model: openai/whisper-large-v3

This model is a W4A8 GPTQ-quantized version of openai/whisper-large-v3. It is intended for automatic speech recognition evaluation and inference on Arm CPU systems using vLLM.

The quantization workflow uses llmcompressor to quantize Whisper linear layers with 4-bit integer weights and 8-bit integer input activations. The language-model projection head is left unquantized. The model was quantized without additional fine-tuning.

Intended Use

This model is intended for:

  • Automatic speech recognition with Whisper-compatible tooling.
  • Experiments comparing BF16 and compressed Whisper inference.
  • Arm CPU vLLM deployments where reduced model size is useful.

This model is not intended to improve transcription quality over the base model. Users should validate quality, latency, memory use, and supported runtime behavior on their target hardware and workload before deployment.

Deployment

Use with vLLM

from vllm import LLM, SamplingParams
from vllm.assets.audio import AudioAsset

llm = LLM(
    model="Arm/whisper-large-v3-quantized.w4a8",
    max_model_len=448,
    max_num_seqs=400,
    limit_mm_per_prompt={"audio": 1},
)

inputs = {
    "encoder_prompt": {
        "prompt": "",
        "multi_modal_data": {
            "audio": AudioAsset("winning_call").audio_and_sample_rate,
        },
    },
    "decoder_prompt": "<|startoftranscript|>",
}

outputs = llm.generate(inputs, SamplingParams(temperature=0.0, max_tokens=128))
print(outputs[0].outputs[0].text)

Quantization Details

Field Value
Base model openai/whisper-large-v3
Quantization method GPTQ
Weight precision INT4
Weight strategy Per-channel, symmetric
Input activation precision INT8
Activation strategy Dynamic, per-token, symmetric
Quantized modules Linear layers
Unquantized modules proj_out
Calibration dataset MLCommons/peoples_speech, subset test, split test
Calibration task prefix English transcription
Calibration samples used for reported run 1,024
Maximum calibration sequence length 2,048

Evaluation

Evaluation was run with lmms-eval using the whisper_vllm model interface on LibriSpeech and FLEURS. Lower WER is better.

Benchmark Split BF16 WER W4A8 WER BF16/W4A8 Recovery
LibiriSpeech (WER) test-clean 2.1517 2.1989 97.9%
LibiriSpeech (WER) test-other 3.9352 4.0865 96.3%
Fleurs (WER) cmn_hans_cn 7.7907 8.1741 95.3%
Fleurs (WER) en 4.0442 4.0785 99.2%

On average our INT4 implementation is able to recover 97.2% of BF16 WER.

Reproduce Quantization

Create a fresh quantization environment:

python -m venv .quantize
source .quantize/bin/activate
pip install llmcompressor
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu

Run quantization:

python quantize.py \
  --model_path openai/whisper-large-v3 \
  --save_dir data/model_dir \
  --num_calibration_samples 1024

The output is written to:

data/model_dir/whisper-large-v3-quantized.w4a8

Reproduce Evaluation

Create a fresh evaluation environment:

python -m venv .eval
source .eval/bin/activate
export VLLM_VERSION=0.23.0
pip install "https://github.com/vllm-project/vllm/releases/download/v${VLLM_VERSION}/vllm-${VLLM_VERSION}+cpu-cp38-abi3-manylinux_2_34_aarch64.whl" --extra-index-url https://download.pytorch.org/whl/cpu
pip install editdistance
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu

git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git ~/lmms-eval
pip install -e ~/lmms-eval

Run the W4A8 evaluation:

lmms-eval \
  --model whisper_vllm \
  --model_args "pretrained=Arm/whisper-large-v3-quantized.w4a8" \
  --tasks librispeech_test_other,librispeech_test_clean,fleurs \
  --batch_size 64 \
  --output_path results/w4a8/full_suite

Limitations

  • Reported evaluations cover LibriSpeech and two FLEURS language splits only.
  • The calibration samples use English transcription examples.
  • Runtime support depends on vLLM, compressed-tensors, and target hardware.
  • Quantization can change outputs, especially on languages, accents, domains, audio conditions, and decoding settings not covered by the reported evaluation.

Ethical Considerations

This model inherits the capabilities and limitations of Whisper large-v3. ASR systems can produce incorrect transcripts and may perform unevenly across languages, accents, dialects, speakers, domains, and recording conditions. Do not use transcripts as the sole basis for high-stakes decisions without human review.

About this version

This repository contains a W4A8 quantized version of OpenAI’s Whisper large-v3 model. Arm quantized the model using INT4 weight quantization and INT8 activation quantization to enable more efficient execution with vLLM on Arm-based platforms. No additional training or fine-tuning was applied by Arm. The original model architecture, intended automatic speech recognition and speech translation use cases, and known limitations remain applicable, although quantization may affect numerical behavior and accuracy.

Original model and documentation

For full details of the original model, please refer to the original OpenAI Whisper large-v3 model card: https://huggingface.co/openai/whisper-large-v3

Purpose of this release

Arm provides this quantized model to enable developers to evaluate and build applications using W4A8 Whisper large-v3 inference with vLLM on Arm-based systems. Users should validate its accuracy and behavior under the languages, audio conditions, decoding settings, and deployment environment relevant to their application.

Downloads last month
48
Safetensors
Model size
2B params
Tensor type
F32
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Arm/whisper-large-v3-quantized.w4a8

Quantized
(47)
this model

Dataset used to train Arm/whisper-large-v3-quantized.w4a8