Downloads · 30 days
174
2% of all-time downloads
spicyneuron/GLM-5.2-MLX-4.5bit
GLM-5.2-MLX-4.5bit is a text generation model from spicyneuron. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
zai-org/GLM-5.2 optimized for running on a Mac Studio M3 512.
Downloads · 30 days
174
2% of all-time downloads
All-time downloads
8.2K
Public
Parameters
743B
421 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors421 GB · 100%
How the weights are stored.
U32743B · 100%
From the Hugging Face model README
zai-org/GLM-5.2 optimized for running on a Mac Studio M3 512.
NOTE: Run with https://github.com/ml-explore/mlx-lm/pull/1410 until the PR is merged.
# Start server at http://localhost:8080/v1/chat/completions
uvx --from mlx-lm mlx_lm.server \
--host 127.0.0.1 \
--port 8080 \
--model spicyneuron/GLM-5.2-MLX-4.5bit
| metric | this model |
|---|---|
| bpw | 4.535 |
| base memory | 392.454 |
| peak memory (1024/512) | 422.787 |
| prompt tok/s (1024) | 194.114 ± 0.079 |
| gen tok/s (512) | 17.781 ± 0.028 |
| kl mean* | 0.049 ± 0.002 |
| kl p95 | 0.113 ± 0.002 |
| perplexity | 4.642 ± 0.036 |
| arc_challenge | 0.690 ± 0.021 |
| hellaswag | 0.780 ± 0.019 |
* KL calculated against the largest quant I could run locally (~5.3 bit). Real KL is against FP will be higher.
Quantized with a mlx-lm fork. MLX quantization options differ than llama.cpp, but the principles are the same: