File size: 2,757 Bytes
e3e3f8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2b42915
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e3e3f8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
---
license: mit
base_model: zai-org/GLM-5.2
base_model_relation: quantized
pipeline_tag: text-generation
library_name: mlx
tags:
- mlx
- moe
- glm
- text-generation
---

# GLM-5.2-MLX-4bit

## Runtime — updated 2026-08-28: load with `--trust-remote-code`

This repository now bundles `glm_moe_dsa.py` (declared via `model_file` in `config.json`), a fixed runtime
for this architecture, and needs it:

```bash
mlx_lm.generate --model pipenetwork/GLM-5.2-MLX-4bit --trust-remote-code --prompt "..." --max-tokens 300
```

mlx-lm's own `glm_moe_dsa` builds a lightning indexer on all 78 layers, but GLM-5.2 ships indexer weights
on 21 (`indexer_types`: the other 57 "shared" layers reuse the previous full layer's top-k selection).
`mlx_lm.load` loads leniently and left those 57 indexers at random initialisation. Prompts up to 2048
tokens were unaffected (the indexer is bypassed below `index_topk`); beyond that, 57 of 78 layers attended
to keys chosen by random projections. The bundled runtime implements the schedule as the reference does
(plus fp32 indexer scores and router logits and the indexer LayerNorm epsilon); tiny-config parity against
`transformers` 5.16 is 4e-7 with the sparse path live, and a strict load of this checkpoint reports zero
missing and zero unexpected tensors. Details, tests and the GLM-5.3 builds made with it:
[github.com/PipeNetwork/glm53-mlx](https://github.com/PipeNetwork/glm53-mlx). The weights are unchanged.


MLX (Apple Silicon) conversion of [zai-org/GLM-5.2](https://huggingface.co/zai-org/GLM-5.2) — a `glm_moe_dsa` MoE (256 experts, DeepSeek-V3.2-style sparse attention) — quantized to **4-bit**.

## Quantizations
Part of the [**GLM-5.2 MLX** collection](https://huggingface.co/collections/pipenetwork/glm-52-mlx-6a31fa56e37a8ac73daf25b7).

| Variant | Notes |
|---|---|
| [8-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-8bit) | 8-bit · ~800GB · needs ~1TB RAM · integrity-checked |
| [6-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-6bit) | 6-bit · ~625GB · needs ~768GB RAM · integrity-checked |
| [5-bit](https://huggingface.co/pipenetwork/GLM-5.2-MLX-5bit) | 5-bit · ~530GB · needs ~640GB RAM · integrity-checked |
| **4-bit** (this repo) | 4-bit · ~430GB · tight on 512GB · smoke-tested |
| [mixed](https://huggingface.co/pipenetwork/GLM-5.2-MLX-mixed-3_6bit) | mixed · experts@3-bit / non-expert@6-bit · ~360GB · 512GB-fit · smoke-tested |

## Use with mlx-lm
```bash
pip install mlx-lm
python -m mlx_lm generate --model pipenetwork/GLM-5.2-MLX-4bit --prompt "Hello" -m 256
```

## Validation
Smoke-tested locally (loads + generates coherent text).

## License
MIT (inherited from base). Quantization config (excerpt): `{"group_size": 64, "bits": 4, "mode": "affine"}`.