--- license: mit base_model: zai-org/GLM-5.2 base_model_relation: quantized pipeline_tag: text-generation library_name: mlx tags: - mlx - moe - glm - text-generation --- # GLM-5.2-MLX-4bit ## Runtime — updated 2026-08-28: load with `--trust-remote-code` This repository now bundles `glm_moe_dsa.py` (declared via `model_file` in `config.json`), a fixed runtime for this architecture, and needs it: ```bash mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-4bit --trust-remote-code --prompt "..." --max-tokens 300 ``` mlx-lm's own `glm_moe_dsa` builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights on 21 (`indexer_types`: the other 57 "shared" layers reuse the previous full layer's top-k selection). `mlx_lm.load` loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048 tokens were unaffected (the indexer is bypassed below `index_topk`); beyond that, 57 of 78 layers attended to keys chosen by random projections. The bundled runtime implements the schedule as the reference does (plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against `transformers` 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it: [github.com/PipeNetwork/glm53-mlx](https://github.com/PipeNetwork/glm53-mlx). The weights are unchanged. MLX (Apple Silicon) conversion of [zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2) — a `glm_moe_dsa` MoE (256 experts, DeepSeek-V3.2-style sparse attention) — quantized to **4-bit**. ## Quantizations Part of the [**GLM-5.2 MLX** collection](https://huggingface.co/collections/pipenetwork/glm-52-mlx-6a31fa56e37a8ac73daf25b7). | Variant | Notes | |---|---| | [8-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-8bit) | 8-bit · ~800GB · needs ~1TB RAM · integrity-checked | | [6-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-6bit) | 6-bit · ~625GB · needs ~768GB RAM · integrity-checked | | [5-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-5bit) | 5-bit · ~530GB · needs ~640GB RAM · integrity-checked | | **4-bit** (this repo) | 4-bit · ~430GB · tight on 512GB · smoke-tested | | [mixed](https://huggingface.co/pipenetwork/GLM-5.2-MLX-mixed-3_6bit) | mixed · experts@3-bit / non-expert@6-bit · ~360GB · 512GB-fit · smoke-tested | ## Use with mlx-lm ```bash pip install mlx-lm python -m mlx_lm generate --model pipenetwork/GLM-5.2-MLX-4bit --prompt "Hello" -m 256 ``` ## Validation Smoke-tested locally (loads + generates coherent text). ## License MIT (inherited from base). Quantization config (excerpt): `{"group_size": 64, "bits": 4, "mode": "affine"}`.