Qwen3-TTS 1.7B Base — CoreML, 1024-position stateful decoder

Experimental six-component CoreML export of Qwen3-TTS 1.7B Base. This repository includes the conversion source, a Python reference runner, and numerical validation against the original checkpoint. The CodeDecoder uses MLState with 1024 cache positions. FP32 decoder computation avoids the FP16 overflow observed during checkpoint validation. State tensors remain FP16. MultiCodeDecoder includes the trained 2048-to-1024 input projection and uses its own 16-position explicit cache.

The speech-swift CoreML runtime supports this 1.7B/1024 bundle from PR #489, alongside the existing 0.6B default. It reads bundle dimensions and requires a prepared 2048-channel speaker embedding. CPU is the default execution route for this bundle. The included Python runner remains available.

Property Value
Upstream model Qwen3-TTS-12Hz-1.7B-Base
Upstream revision fd4b254389122332181a7c3db7f27e918eec64e3
Format Six compiled .mlmodelc components
Talker / speaker dimensions 2048
Code predictor dimensions 1024, with trained input projection
CodeDecoder capacity 1024 positions, including prompt and audio steps
SpeechDecoder capacity 125 frames per call, 10 seconds
Audio Mono 24 kHz; 1920 samples/frame (12.5 frames/s)
Quantization No weight palettization; FP32 CodeDecoder and MultiCodeDecoder, FP16 other components
Platform Apple Silicon, macOS 15+; compiled with Core ML on macOS 26.6.2

Files

Component Size (GB) Computation
CodeDecoder.mlmodelc 5.664 FP32
CodeEmbedder.mlmodelc 0.013 FP16
MultiCodeDecoder.mlmodelc 0.449 FP32
MultiCodeEmbedder.mlmodelc 0.126 FP16
SpeechDecoder.mlmodelc 0.228 FP16
TextProjector.mlmodelc 0.639 FP16
Total compiled models 7.119

source/ contains the exporter, pinned dependencies, reference runner, regression tests, and real-checkpoint validation harness. validation/ contains machine-readable numerical results. Tokenizer vocabulary and special text embeddings are included. Supply a 2048-dimensional speaker embedding from your own reference recording; the 0.6B model's speaker embedding is incompatible.

Reproduce conversion

See the full conversion guide. Run components sequentially in separate processes to limit memory use.

python3.11 -m venv .venv
source .venv/bin/activate
pip install -r source/requirements-coreml.txt
python source/convert_coreml.py \
  --model-id Qwen/Qwen3-TTS-12Hz-1.7B-Base \
  --revision fd4b254389122332181a7c3db7f27e918eec64e3 \
  --tokenizer-revision 7dd38ad4e9bad454aae9cd937d0cd577604fe229 \
  --max-seq-len 1024 --only CodeDecoder --compile --output-dir bundle

The guide lists all six components and the Embeddings preparation step. The exporter preserves normalized talker hidden states, the code predictor's input projection, the vocoder's sliding attention window, and single writes to each state buffer.

Run inference

Download the compiled bundle, then prepare a speaker embedding from a reference recording you are authorized to use. The Embeddings step downloads the upstream checkpoint for speaker extraction; it does not repeat CoreML conversion.

hf download aufklarer/Qwen3-TTS-1.7B-CoreML --local-dir bundle
python bundle/source/convert_coreml.py \
  --model-id Qwen/Qwen3-TTS-12Hz-1.7B-Base \
  --revision fd4b254389122332181a7c3db7f27e918eec64e3 \
  --tokenizer-revision 7dd38ad4e9bad454aae9cd937d0cd577604fe229 \
  --max-seq-len 1024 --only Embeddings \
  --reference-audio reference.wav --output-dir bundle
python bundle/source/run_coreml.py bundle "Hello, how are you today?" \
  --speaker-embedding bundle/speaker_embedding.npy --compute cpu --output hello.wav

With the updated speech-swift runtime, use the same prepared bundle:

speech qwen3-tts-coreml "Hello, how are you today?" \
  --model-directory bundle \
  --speaker-embedding bundle/speaker_embedding.npy --output hello.wav

Or load the published model from Swift:

import Foundation
import Qwen3TTSCoreML

let model = try await Qwen3TTSCoreMLModel.fromPretrained(
    modelId: Qwen3TTSCoreMLModel.largeModelId,
    speakerEmbeddingURL: URL(fileURLWithPath: "bundle/speaker_embedding.npy")
)
let audio = try model.synthesize(text: "Hello, how are you today?")

See the Swift inference guide for configuration and CLI options.

Or use the Python interface:

import sys
import numpy as np
sys.path.insert(0, "bundle/source")
from run_coreml import Pipeline
pipeline = Pipeline("bundle", compute="cpu")
audio, codes, metrics = pipeline.synthesize(
    "Hello, how are you today?",
    np.load("bundle/speaker_embedding.npy"),
    seed=42,
)

Each request creates a fresh talker state. The runner checks cache and SpeechDecoder capacity rather than silently truncating waveforms. EOS status and whether a generation limit was reached should be checked for longer text. A 1024-position cache does not make SpeechDecoder longer: it remains a separate 125-frame graph. Re-export and validate a larger --speech-frames value if longer single-call audio is required.

Validation and performance

All six compiled components pass CPU checks against the original PyTorch checkpoint (minimum cosine similarity across outputs: 0.999306). The CodeDecoder minimum is 0.99999988, including positions 255, 256, 1022, and 1023. The full numerical reports are in validation/*-cpu.json.

A separate test completed 1024 consecutive state updates with seeded synthetic embeddings: every cache slot was written, all outputs and states were finite, the first slot remained intact, and fresh-state reset was exact. This tests sustained cache operation, not 80 seconds of continuous speech.

Two end-to-end English samples reached EOS and transcribed exactly with Parakeet-TDT v3 on CPU (0% WER for these two sentences). The reference speaker embedding used for these checks was prepared from the synthetic Kokoro regression fixture in speech-swift; it is not included as a default voice.

Timings below are single CPU smoke runs on an Apple M5 Pro (48 GB, macOS 26.6.2) while another workload was active, not a controlled performance comparison. Model loading took 6.41 seconds and is excluded from generation times. Generation timing starts after initial text tokenization/projection and includes talker prefill, audio-code generation, and waveform decoding. RTF is wall time divided by audio duration; lower is faster.

Text Audio seconds Generation seconds RTF
Hello, how are you today? 2.24 28.36 12.66
The quick brown fox jumps over the lazy dog. 2.96 25.26 8.53

An expanded 12-sentence CPU run reached EOS on every sentence and scored 0% WER across 92 words. Median full-generation latency was 26.07 s, p95 76.80 s, and aggregate RTF 14.53 (430.01 s to generate 29.60 s of audio). Loading took 12.85 s; peak sampled synthesis process-tree RSS was 6.70 GiB. These measurements include the first request and exclude loading. Another speech workload remained active, so the timings are not controlled performance comparisons. Twenty focused regression tests pass. See the benchmark report and per-sentence results for methods and scope.

Swift runtime validation

Sixteen focused unit/CLI tests and all twelve targeted CoreML E2E tests passed with no skips. The 1.7B suite covers a prompt beyond 256 positions, exact request-state reset, invalid frame/speaker inputs, and real synthesis followed by Qwen3-ASR transcription. The reference sentence, “The quick brown fox jumps over the lazy dog.”, transcribed exactly. The CLI produced byte-identical audio to the API test. A deterministic 0.6B fixture also remained byte-identical to the previous runtime; its ASR transcription omits the final word in both versions.

A debug CPU CLI smoke run generated 3.12 seconds of audio in 14.548 seconds (excluding loading). This is not a release benchmark or a controlled comparison with the Python measurements above. The isolated E2E run peaked at 10.27 GiB sampled process-tree RSS, including ASR. See the Swift validation summary. GPU/ANE routing and iOS execution for 1.7B remain unvalidated.

The FP16 CodeDecoder was rejected after it produced a zero output at position 256 and failed parity at later positions. This bundle uses FP32 computation for CodeDecoder and MultiCodeDecoder; caches and external embeddings are FP16.

CPU execution is the release validation route. GPU, Neural Engine placement, iOS devices, and long continuous synthesis are not validated by these results. The 1024-step state test uses synthetic embeddings. Long continuous speech generation and larger SpeechDecoder frame capacities require separate checks.

Sources and links

Downloads last month
43
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aufklarer/Qwen3-TTS-1.7B-CoreML

Finetuned
(44)
this model

Collection including aufklarer/Qwen3-TTS-1.7B-CoreML