You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Apertus-v1.5-70B-FP8-mlx

8-bit MLX quantization of the full multimodal swiss-ai/Apertus-v1.5-70B, with the vision and audio towers intact.

Images, audio and text in. Text out.

This model reads images and audio; it does not produce them. The embedding table covers 266,752 tokens but the output head covers only the first 131,072 (output_vocab_size), so image and audio codes can only ever be inputs. There is no image decoder and no audio decoder in this checkpoint.

How the multimodality works

Apertus 1.5 does not use a projector or cross-attention. Its vision and audio towers are tokenizers: they emit codebook indices, a fixed offset is added, and the resulting ids are embedded by the same embed_tokens that handles text. Fusion is token-id splicing.

range contents
0 – 131,071 text
131,072 – 131,271 200 reserved OMNI control tokens
131,272 – 262,343 image codes (image_token_offset = 131,272)
262,344 – 266,439 audio codes (audio_token_offset = 262,344)
266,440 – 266,751 padding to a round embedding size

The consequence for quantization is the important part. Code assignment is an argmax over a 131,072-entry visual codebook and a Euclidean-argmin over a 4,096-entry audio codebook. A half-precision cast flips code ids that are otherwise deterministic, and a wrong code is not a small error — it is a different visual token. So both towers are kept in float32 and unquantized in this build, while the language model is quantized to 8 bits.

That is not a hand-applied exception: the mlx-vlm implementation declares it on the model itself, through cast_predicate (excludes the towers from the bfloat16 cast) and quant_predicate (excludes them from quantization). A plain mlx_vlm.convert -q does the right thing with no extra flags.

Model details

Parameters 70,599,864,640 (language model) + 228,540,000 (towers)
Quantization 8-bit affine, group size 64 (8.575 bits/weight) — language model only
Towers float32, unquantized, 914 MB
Size on disk 77.1 GB / 71.8 GiB (15 shards, 77,108,760,640 bytes)
Input vocabulary 266,752
Output vocabulary 131,072 (text only)
Context length 262,144
Source precision bfloat16 (language model), float32 (towers)
Source revision 59e744e
Converted with mlx-vlm at 6c739c1, mlx 0.32.2, transformers 5.17.0

Requirements

This needs the unmerged mlx-vlm pull request that adds apertus1p5. Released mlx-vlm does not have it, and mlx-lm cannot run this model at all.

git clone https://github.com/Gusanidas/mlx-vlm.git
cd mlx-vlm && git checkout 6c739c192407e48ff6520703663320c3c506f52c
uv venv --python 3.12 .venv && uv pip install -e .

Blaizzy/mlx-vlm#1847 is open, not merged. The Apertus maintainers asked to hold it until huggingface/transformers#47662 lands, because the checkpoint layout and the reference implementation are not final. Pin the commit above. If the upstream layout changes, this build may need reconverting.

The branch does not merge cleanly onto current mlx-vlm main. Use the branch's own tree; do not rebase it.

Usage

# an image
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
  --image photo.jpg --prompt "Describe this image." \
  --max-tokens 300 --temperature 0.0

# audio -- 24 kHz mono only, see below
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
  --audio speech.wav --prompt "Transcribe this." --max-tokens 200

# both at once
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
  --image photo.jpg --audio speech.wav \
  --prompt "What is in the image, and what does the speaker say?"
from mlx_vlm.utils import load
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm import generate

model, processor = load("./Apertus-v1.5-70B-FP8-mlx")
messages = [{"role": "user", "content": "Describe this image."}]
prompt = apply_chat_template(processor, model.config, messages, num_images=1)
print(generate(model, processor, prompt, image=["photo.jpg"], max_tokens=300))

Audio must be 24 kHz mono

The processor raises on anything else. It does not resample and does not downmix.

ffmpeg -i input.m4a -ac 1 -ar 24000 -f wav output.wav

Token budgeting

Media is expensive in context, and the cost is predictable:

input codes
image, minimum (min_pixels, 256x256) 256
image, maximum (max_pixels, 1400x1400 -> 88x88 grid) 7,744
audio 40 codes per second

An image is downsampled 16x spatially, so the code count is (height // 16) * (width // 16) after the processor's resize.

Thinking mode

Apertus 1.5 reasons between <|inner_prefix|> (id 32) and <|inner_suffix|> (id 33) — not between <think> and </think>, which it reserves at ids 69 and 70 but never emits. The chat template gates on enable_thinking and renders Deliberation: enabled or disabled in the developer block; mlx-vlm defaults it to off.

mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx --enable-thinking \
  --thinking-start-token "<|inner_prefix|>" --thinking-end-token "<|inner_suffix|>" \
  --prompt "..." --max-tokens 800

For the OpenAI-compatible server, pass the same two flags — or set MLX_VLM_THINKING_START_TOKEN / MLX_VLM_THINKING_END_TOKEN:

mlx_vlm.server --model ./Apertus-v1.5-70B-FP8-mlx --enable-thinking \
  --thinking-start-token "<|inner_prefix|>" --thinking-end-token "<|inner_suffix|>"

Without those flags the server looks for </think>, never finds it, and the entire reasoning block lands in content with a literal <|inner_prefix|> in front of it while reasoning_content stays null. Verified on this build.

Budget generously for --max-tokens: the reasoning span is often several times longer than the answer, and a truncated generation can stop before the answer starts.

Memory

text (8k KV cache) with one image
peak 77.5 GB 83.3 GB

Measured peaks on this build. The 262,144-token context will exhaust memory long before the weights do — a 66k-token prompt at 8-bit peaks near 101 GB on a 128 GB machine, which is close enough to the default Metal wired limit (~75% of RAM) to matter. Cap with --max-kv-size, quantize the cache with --kv-bits 8, or raise iogpu.wired_limit_mb.

Conversion recipe

mlx_vlm.convert --hf-path <clean copy of the checkout> \
  --mlx-path ./Apertus-v1.5-70B-FP8-mlx \
  -q --q-bits 8 --q-group-size 64

Peak host memory during conversion was 17 GB against the 144 GB bf16 source: weights are read lazily through mmap, and save_weights donates and evaluates one shard at a time.

Convert from a directory that contains no subdirectories. After writing the weights, mlx_vlm/convert.py runs shutil.copytree over every entry of model_path.iterdir() that is a directory — and iterdir() includes dotfiles. Pointed at a git clone of the source repo, it will spend hours copying .git/lfs into the output. Hardlink the safetensors and copy the small JSON files into a staging directory instead.

What sanitize() changes and drops

The towers are the upstream float32 weights, rearranged for MLX — not retrained and not requantized. Of the 266 tower tensors shared with the source checkpoint, 193 are bit-identical and 73 differ only by documented layout transforms: 72 convolution weights transposed from NCHW to NHWC, plus weight-norm folding (w = g * v / ||v||) and LSTM gate-bias folding (b_ih + b_hh) in the audio encoder.

207 upstream tensors are dropped, all of them unreachable:

  • 158 audio_tokenizer.backbone.* and 2 audio_tokenizer.head.* — the Vocos decoder. The language model cannot emit audio codes, so nothing can ever call it.
  • 44 audio_tokenizer.encoder.* — the weight_g/weight_v pairs consumed by weight-norm folding.
  • 3 audio_tokenizer.quantizer.* — EMA training buffers.

Verification

Everything below was measured on this build.

Structural — every one of the 290 tower tensors is float32 with no .scales/.biases sibling; the language model is quantized (482 tensors carry scales); embed_tokens has 266,752 rows and lm_head has 131,072; the quantization block lists no tower module; tokenizer_config.json keeps its omnimodal_config offsets (131,272 / 262,344), processor_class and both tokenizer blocks.

Format tokens — all 14 core and assistant tokens from the Apertus format spec resolve to the documented ids (system 61/62, developer 63/64, user 65/66, assistant 67/68, inner 32/33, tools 71/72, tool output 73/74), and eos_token_id is [2, 68, 72]. A rendered three-turn prompt contains exactly one BOS.

Functional — correct description of a photograph including its discriminating details; exact transcription of a 5-second speech clip; interleaved image + audio in one prompt; multi-turn context retained; thinking mode emitting 32/33 and never 69/70; deterministic vision codes across repeated encodes; the documented ValueError on batch size > 1; and the processor correctly rejecting non-24 kHz and non-mono audio.

The 6-bit and 8-bit builds produce bit-identical image and audio codes. Their 290 tower tensors are bit-identical to each other, which is what keeping the towers out of quantization is for.

Evaluation

The language model in this build is bit-identical to the text-only 8-bit conversion that preceded it and that this build replaces. The key sets match exactly, 30 of 30 randomly sampled tensors across 23 layers compare bit-for-bit equal, and the first 131,072 rows of embed_tokens (weights and scales) are unchanged — the new table only appends the media rows. So on text-only input this build is numerically identical to its predecessor, and the perplexity and top-1 agreement numbers measured there transfer exactly.

Perplexity on allenai/tulu-3-sft-mixture, 512 sequences of 512 tokens (261,632 scored tokens), seed 123, both quantizations scored on byte-identical batches:

6-bit 8-bit
Perplexity 6.7565 6.7341
Size on disk 59.2 GB 77.1 GB
Peak memory, text (8k KV cache) 59.5 GB 77.5 GB
Peak memory, one image 65.4 GB 83.3 GB
Generation 9.29 tok/s 7.28 tok/s

6-bit costs +0.333% perplexity relative to 8-bit. The paired difference in mean token loss is +0.003325 +/- 0.001057 nats.

Top-1 agreement

Perplexity measures loss, not behaviour. The behavioural question is how often the two builds would actually pick a different token. Measured over the same 261,632 positions:

Same argmax 94.386% (95% CI [94.162%, 94.611%])
Different argmax 5.614% (14,687 positions)
Next-token accuracy 6-bit 64.684% vs 8-bit 64.767%
Paired accuracy difference +0.083 +/- 0.029 pp (t = 2.85)

The two models pick a different token roughly once every 18 positions -- but disagreements cluster almost entirely on near-ties. Where the models agree, the median top1-top2 logit margin is 3.000; where they disagree it is 0.250, and 92.1% of disagreements fall below a 1.0 margin. The builds diverge exactly where the model is close to indifferent between two candidates.

And when they diverge, neither is meaningfully more correct. Of the 14,687 disagreements, both models are wrong 60.71% of the time; 8-bit is right on 20.39% and 6-bit on 18.90%. Among the cases actually decided, 8-bit wins 51.9% -- barely distinguishable from a coin flip.

Two different questions, two different answers. Will the output differ? Yes, visibly -- at 1-in-18 divergence, compounding autoregressively, a 200-token answer will read differently between the two builds. Do not treat them as interchangeable if you need reproducible output. Will it be worse? Essentially no: 0.083 pp of accuracy.

None of this covers the vision or audio path. No multimodal benchmark was run. What was checked is that both builds emit identical media codes and that both describe a photograph and transcribe speech correctly — a qualitative check, not a benchmark. Upstream benchmark figures do not transfer either: they were measured on the full-precision model.

Caveats

  • 6-bit is compared against 8-bit, not against the bf16 original, which does not fit in 128 GB. 8-bit is normally very close to bf16, so it serves as a reasonable proxy reference, but 0.333% is a relative figure.
  • Divergence compounds during free generation in a way teacher forcing does not capture.
  • Scored under teacher forcing on reference text, not on the models' own generations.
  • The mlx-vlm implementation's own correctness numbers (744/744 and 418/418 code exactness against the reference implementation) are the PR author's, measured on the 8B model, and were not reproduced here. The tower-dtype assertions and determinism checks above are a proxy for them, not a substitute.

Which to use

This 8-bit build is the higher-fidelity option for text and the right choice when quality matters more than speed and throughput.

Note that it is not sharper on images or audio: the towers are float32 and unquantized in both builds, and the two produce bit-identical media codes. The difference is confined to the language model.

Prefer the 6-bit build if you want ~28% more throughput and ~18 GB more headroom — which matters here, since one image already pushes this build to 83 GB. It costs 0.333% perplexity and 0.083 pp of next-token accuracy, neither likely to be visible in practice.

Measured on

Apple M5 Max MacBook Pro, 128 GB unified memory, macOS 26.6.2, Python 3.12, mlx 0.32.2, transformers 5.17.0. Throughput is single-stream, greedy decoding.

Limitations

Inherits all limitations of the base model: generated content may not be factually accurate, logically consistent, or free from biases in the training data. Use as an assistive tool, not a definitive source, and verify important information. See the Apertus Charter for the alignment principles governing the model's responses.

For this build specifically:

  • Batch size 1 only. Anything else raises ValueError; it is a hard guard in the model, not a silent fallback.
  • Audio must be 24 kHz mono. The processor raises rather than resampling.
  • Text out only — no image or audio generation.
  • Quantization introduces the small measured quality cost described above.
  • Requires an unmerged mlx-vlm branch, which may change.
  • No output filter is provided.

License and legal

Apache 2.0, inherited from the base model. The full text is in LICENSE.txt, copied unchanged from the upstream repository (it is the stock Apache 2.0 text, with no added copyright notice). Use is also subject to the Apertus Acceptable Use Policy. The upstream repository is gated; access conditions apply to the original weights from which this conversion derives.

EU AI Act transparency documentation:

For removal requests of personally identifiable information or copyrighted content, contact the dataset owners or the Apertus team directly: llm-privacy-requests@swiss-ai.org, llm-copyright-requests@swiss-ai.org.

Credits

Model by the Apertus / Swiss AI Initiative (EPFL, ETH Zurich, CSCS). The MLX implementation of the apertus1p5 towers and processor is by Gusanidas in mlx-vlm#1847. This repository only contains a format conversion and quantization — no additional training.

Contact for the upstream model: https://apertus-ai.org/contact/.

Downloads last month
2
Safetensors
Model size
72B params
Tensor type
F32
·
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sbaechler/Apertus-v1.5-70B-FP8-mlx

Quantized
(7)
this model