Instructions to use sbaechler/Apertus-v1.5-70B-FP8-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sbaechler/Apertus-v1.5-70B-FP8-mlx with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sbaechler/Apertus-v1.5-70B-FP8-mlx") config = load_config("sbaechler/Apertus-v1.5-70B-FP8-mlx") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use sbaechler/Apertus-v1.5-70B-FP8-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sbaechler/Apertus-v1.5-70B-FP8-mlx"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sbaechler/Apertus-v1.5-70B-FP8-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use sbaechler/Apertus-v1.5-70B-FP8-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sbaechler/Apertus-v1.5-70B-FP8-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sbaechler/Apertus-v1.5-70B-FP8-mlx
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sbaechler/Apertus-v1.5-70B-FP8-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sbaechler/Apertus-v1.5-70B-FP8-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sbaechler/Apertus-v1.5-70B-FP8-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Apertus-v1.5-70B-FP8-mlx
8-bit MLX quantization of the full multimodal swiss-ai/Apertus-v1.5-70B, with the vision and audio towers intact.
Images, audio and text in. Text out.
This model reads images and audio; it does not produce them. The embedding table covers 266,752 tokens but the output head covers only the first 131,072 (
output_vocab_size), so image and audio codes can only ever be inputs. There is no image decoder and no audio decoder in this checkpoint.
How the multimodality works
Apertus 1.5 does not use a projector or cross-attention. Its vision and audio
towers are tokenizers: they emit codebook indices, a fixed offset is added,
and the resulting ids are embedded by the same embed_tokens that handles text.
Fusion is token-id splicing.
| range | contents |
|---|---|
| 0 – 131,071 | text |
| 131,072 – 131,271 | 200 reserved OMNI control tokens |
| 131,272 – 262,343 | image codes (image_token_offset = 131,272) |
| 262,344 – 266,439 | audio codes (audio_token_offset = 262,344) |
| 266,440 – 266,751 | padding to a round embedding size |
The consequence for quantization is the important part. Code assignment is an argmax over a 131,072-entry visual codebook and a Euclidean-argmin over a 4,096-entry audio codebook. A half-precision cast flips code ids that are otherwise deterministic, and a wrong code is not a small error — it is a different visual token. So both towers are kept in float32 and unquantized in this build, while the language model is quantized to 8 bits.
That is not a hand-applied exception: the mlx-vlm implementation declares it on
the model itself, through cast_predicate (excludes the towers from the
bfloat16 cast) and quant_predicate (excludes them from quantization). A plain
mlx_vlm.convert -q does the right thing with no extra flags.
Model details
| Parameters | 70,599,864,640 (language model) + 228,540,000 (towers) |
| Quantization | 8-bit affine, group size 64 (8.575 bits/weight) — language model only |
| Towers | float32, unquantized, 914 MB |
| Size on disk | 77.1 GB / 71.8 GiB (15 shards, 77,108,760,640 bytes) |
| Input vocabulary | 266,752 |
| Output vocabulary | 131,072 (text only) |
| Context length | 262,144 |
| Source precision | bfloat16 (language model), float32 (towers) |
| Source revision | 59e744e |
| Converted with | mlx-vlm at 6c739c1, mlx 0.32.2, transformers 5.17.0 |
Requirements
This needs the unmerged mlx-vlm pull request that adds apertus1p5. Released
mlx-vlm does not have it, and mlx-lm cannot run this model at all.
git clone https://github.com/Gusanidas/mlx-vlm.git
cd mlx-vlm && git checkout 6c739c192407e48ff6520703663320c3c506f52c
uv venv --python 3.12 .venv && uv pip install -e .
Blaizzy/mlx-vlm#1847 is open, not merged. The Apertus maintainers asked to hold it until huggingface/transformers#47662 lands, because the checkpoint layout and the reference implementation are not final. Pin the commit above. If the upstream layout changes, this build may need reconverting.
The branch does not merge cleanly onto current mlx-vlm main. Use the branch's
own tree; do not rebase it.
Usage
# an image
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
--image photo.jpg --prompt "Describe this image." \
--max-tokens 300 --temperature 0.0
# audio -- 24 kHz mono only, see below
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
--audio speech.wav --prompt "Transcribe this." --max-tokens 200
# both at once
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx \
--image photo.jpg --audio speech.wav \
--prompt "What is in the image, and what does the speaker say?"
from mlx_vlm.utils import load
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm import generate
model, processor = load("./Apertus-v1.5-70B-FP8-mlx")
messages = [{"role": "user", "content": "Describe this image."}]
prompt = apply_chat_template(processor, model.config, messages, num_images=1)
print(generate(model, processor, prompt, image=["photo.jpg"], max_tokens=300))
Audio must be 24 kHz mono
The processor raises on anything else. It does not resample and does not downmix.
ffmpeg -i input.m4a -ac 1 -ar 24000 -f wav output.wav
Token budgeting
Media is expensive in context, and the cost is predictable:
| input | codes |
|---|---|
image, minimum (min_pixels, 256x256) |
256 |
image, maximum (max_pixels, 1400x1400 -> 88x88 grid) |
7,744 |
| audio | 40 codes per second |
An image is downsampled 16x spatially, so the code count is
(height // 16) * (width // 16) after the processor's resize.
Thinking mode
Apertus 1.5 reasons between <|inner_prefix|> (id 32) and <|inner_suffix|>
(id 33) — not between <think> and </think>, which it reserves at ids 69
and 70 but never emits. The chat template gates on enable_thinking and renders
Deliberation: enabled or disabled in the developer block; mlx-vlm defaults
it to off.
mlx_vlm.generate --model ./Apertus-v1.5-70B-FP8-mlx --enable-thinking \
--thinking-start-token "<|inner_prefix|>" --thinking-end-token "<|inner_suffix|>" \
--prompt "..." --max-tokens 800
For the OpenAI-compatible server, pass the same two flags — or set
MLX_VLM_THINKING_START_TOKEN / MLX_VLM_THINKING_END_TOKEN:
mlx_vlm.server --model ./Apertus-v1.5-70B-FP8-mlx --enable-thinking \
--thinking-start-token "<|inner_prefix|>" --thinking-end-token "<|inner_suffix|>"
Without those flags the server looks for
</think>, never finds it, and the entire reasoning block lands incontentwith a literal<|inner_prefix|>in front of it whilereasoning_contentstaysnull. Verified on this build.
Budget generously for --max-tokens: the reasoning span is often several times
longer than the answer, and a truncated generation can stop before the answer
starts.
Memory
| text (8k KV cache) | with one image | |
|---|---|---|
| peak | 77.5 GB | 83.3 GB |
Measured peaks on this build. The 262,144-token context will exhaust memory long
before the weights do — a 66k-token prompt at 8-bit peaks near 101 GB on a
128 GB machine, which is close enough to the default Metal wired limit (~75% of
RAM) to matter. Cap with --max-kv-size, quantize the cache with --kv-bits 8,
or raise iogpu.wired_limit_mb.
Conversion recipe
mlx_vlm.convert --hf-path <clean copy of the checkout> \
--mlx-path ./Apertus-v1.5-70B-FP8-mlx \
-q --q-bits 8 --q-group-size 64
Peak host memory during conversion was 17 GB against the 144 GB
bf16 source: weights are read lazily through mmap, and save_weights donates
and evaluates one shard at a time.
Convert from a directory that contains no subdirectories. After writing the weights,
mlx_vlm/convert.pyrunsshutil.copytreeover every entry ofmodel_path.iterdir()that is a directory — anditerdir()includes dotfiles. Pointed at agit cloneof the source repo, it will spend hours copying.git/lfsinto the output. Hardlink the safetensors and copy the small JSON files into a staging directory instead.
What sanitize() changes and drops
The towers are the upstream float32 weights, rearranged for MLX — not retrained
and not requantized. Of the 266 tower tensors shared with the source checkpoint,
193 are bit-identical and 73 differ only by documented layout transforms: 72
convolution weights transposed from NCHW to NHWC, plus weight-norm folding
(w = g * v / ||v||) and LSTM gate-bias folding (b_ih + b_hh) in the audio
encoder.
207 upstream tensors are dropped, all of them unreachable:
- 158
audio_tokenizer.backbone.*and 2audio_tokenizer.head.*— the Vocos decoder. The language model cannot emit audio codes, so nothing can ever call it. - 44
audio_tokenizer.encoder.*— theweight_g/weight_vpairs consumed by weight-norm folding. - 3
audio_tokenizer.quantizer.*— EMA training buffers.
Verification
Everything below was measured on this build.
Structural — every one of the 290 tower tensors is float32 with no
.scales/.biases sibling; the language model is quantized (482 tensors carry
scales); embed_tokens has 266,752 rows and lm_head has 131,072; the
quantization block lists no tower module; tokenizer_config.json keeps its
omnimodal_config offsets (131,272 / 262,344), processor_class and both
tokenizer blocks.
Format tokens — all 14 core and assistant tokens from the
Apertus format spec
resolve to the documented ids (system 61/62, developer 63/64, user 65/66,
assistant 67/68, inner 32/33, tools 71/72, tool output 73/74), and
eos_token_id is [2, 68, 72]. A rendered three-turn prompt contains exactly
one BOS.
Functional — correct description of a photograph including its
discriminating details; exact transcription of a 5-second speech clip;
interleaved image + audio in one prompt; multi-turn context retained; thinking
mode emitting 32/33 and never 69/70; deterministic vision codes across repeated
encodes; the documented ValueError on batch size > 1; and the processor
correctly rejecting non-24 kHz and non-mono audio.
The 6-bit and 8-bit builds produce bit-identical image and audio codes. Their 290 tower tensors are bit-identical to each other, which is what keeping the towers out of quantization is for.
Evaluation
The language model in this build is bit-identical to the text-only 8-bit conversion that preceded it and that this build replaces. The key sets match exactly, 30 of 30 randomly sampled tensors across 23 layers compare bit-for-bit equal, and the first 131,072 rows of
embed_tokens(weights and scales) are unchanged — the new table only appends the media rows. So on text-only input this build is numerically identical to its predecessor, and the perplexity and top-1 agreement numbers measured there transfer exactly.
Perplexity on allenai/tulu-3-sft-mixture, 512 sequences of 512 tokens
(261,632 scored tokens), seed 123, both quantizations scored on byte-identical
batches:
| 6-bit | 8-bit | |
|---|---|---|
| Perplexity | 6.7565 | 6.7341 |
| Size on disk | 59.2 GB | 77.1 GB |
| Peak memory, text (8k KV cache) | 59.5 GB | 77.5 GB |
| Peak memory, one image | 65.4 GB | 83.3 GB |
| Generation | 9.29 tok/s | 7.28 tok/s |
6-bit costs +0.333% perplexity relative to 8-bit. The paired difference in mean token loss is +0.003325 +/- 0.001057 nats.
Top-1 agreement
Perplexity measures loss, not behaviour. The behavioural question is how often the two builds would actually pick a different token. Measured over the same 261,632 positions:
| Same argmax | 94.386% (95% CI [94.162%, 94.611%]) |
| Different argmax | 5.614% (14,687 positions) |
| Next-token accuracy | 6-bit 64.684% vs 8-bit 64.767% |
| Paired accuracy difference | +0.083 +/- 0.029 pp (t = 2.85) |
The two models pick a different token roughly once every 18 positions -- but disagreements cluster almost entirely on near-ties. Where the models agree, the median top1-top2 logit margin is 3.000; where they disagree it is 0.250, and 92.1% of disagreements fall below a 1.0 margin. The builds diverge exactly where the model is close to indifferent between two candidates.
And when they diverge, neither is meaningfully more correct. Of the 14,687 disagreements, both models are wrong 60.71% of the time; 8-bit is right on 20.39% and 6-bit on 18.90%. Among the cases actually decided, 8-bit wins 51.9% -- barely distinguishable from a coin flip.
Two different questions, two different answers. Will the output differ? Yes, visibly -- at 1-in-18 divergence, compounding autoregressively, a 200-token answer will read differently between the two builds. Do not treat them as interchangeable if you need reproducible output. Will it be worse? Essentially no: 0.083 pp of accuracy.
None of this covers the vision or audio path. No multimodal benchmark was run. What was checked is that both builds emit identical media codes and that both describe a photograph and transcribe speech correctly — a qualitative check, not a benchmark. Upstream benchmark figures do not transfer either: they were measured on the full-precision model.
Caveats
- 6-bit is compared against 8-bit, not against the bf16 original, which does not fit in 128 GB. 8-bit is normally very close to bf16, so it serves as a reasonable proxy reference, but 0.333% is a relative figure.
- Divergence compounds during free generation in a way teacher forcing does not capture.
- Scored under teacher forcing on reference text, not on the models' own generations.
- The mlx-vlm implementation's own correctness numbers (744/744 and 418/418 code exactness against the reference implementation) are the PR author's, measured on the 8B model, and were not reproduced here. The tower-dtype assertions and determinism checks above are a proxy for them, not a substitute.
Which to use
This 8-bit build is the higher-fidelity option for text and the right choice when quality matters more than speed and throughput.
Note that it is not sharper on images or audio: the towers are float32 and unquantized in both builds, and the two produce bit-identical media codes. The difference is confined to the language model.
Prefer the 6-bit build if you want ~28% more throughput and ~18 GB more headroom — which matters here, since one image already pushes this build to 83 GB. It costs 0.333% perplexity and 0.083 pp of next-token accuracy, neither likely to be visible in practice.
Measured on
Apple M5 Max MacBook Pro, 128 GB unified memory, macOS 26.6.2, Python 3.12, mlx 0.32.2, transformers 5.17.0. Throughput is single-stream, greedy decoding.
Limitations
Inherits all limitations of the base model: generated content may not be factually accurate, logically consistent, or free from biases in the training data. Use as an assistive tool, not a definitive source, and verify important information. See the Apertus Charter for the alignment principles governing the model's responses.
For this build specifically:
- Batch size 1 only. Anything else raises
ValueError; it is a hard guard in the model, not a silent fallback. - Audio must be 24 kHz mono. The processor raises rather than resampling.
- Text out only — no image or audio generation.
- Quantization introduces the small measured quality cost described above.
- Requires an unmerged mlx-vlm branch, which may change.
- No output filter is provided.
License and legal
Apache 2.0, inherited from the base model. The full text is in LICENSE.txt, copied unchanged from the upstream repository (it is the stock Apache 2.0 text, with no added copyright notice). Use is also subject to the Apertus Acceptable Use Policy. The upstream repository is gated; access conditions apply to the original weights from which this conversion derives.
EU AI Act transparency documentation:
For removal requests of personally identifiable information or copyrighted content, contact the dataset owners or the Apertus team directly: llm-privacy-requests@swiss-ai.org, llm-copyright-requests@swiss-ai.org.
Credits
Model by the Apertus / Swiss AI Initiative (EPFL,
ETH Zurich, CSCS). The MLX implementation of the apertus1p5 towers and
processor is by Gusanidas in
mlx-vlm#1847. This repository
only contains a format conversion and quantization — no additional training.
Contact for the upstream model: https://apertus-ai.org/contact/.
- Downloads last month
- 2
8-bit
Model tree for sbaechler/Apertus-v1.5-70B-FP8-mlx
Base model
swiss-ai/Apertus-v1.5-70B