Downloads · 30 days
150
15% of all-time downloads
samithaj/GLM-4.7-Flash-MTP-4bit
GLM-4.7-Flash-MTP-4bit is a machine learning model from samithaj. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as mit.
The trained multi-token-prediction (MTP / nextn) layer of zai-org/GLM-4.7-Flash, split into a standalone MLX drafter checkpoint and quantized to 4-bit. This is not a standalone language model — it is a single-layer dr…
Downloads · 30 days
150
15% of all-time downloads
All-time downloads
984
Public
Parameters
1.3B
739 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors719 MB · 97%
How the weights are stored.
U321.3B · 100%
From the Hugging Face model README
The trained multi-token-prediction (MTP / nextn) layer of zai-org/GLM-4.7-Flash, split into a standalone MLX drafter checkpoint and quantized to 4-bit. This is not a standalone language model — it is a single-layer draft head that predicts one token ahead from a target model's hidden states, for speculative decoding against GLM-4.7-Flash (pairs with mlx-community/GLM-4.7-Flash-4bit).
GLM-4.7-Flash ships this layer inside the full checkpoint at
model.layers.47.*; MLX conversions of the base model strip it (sanitize
drops layers past num_hidden_layers), so quantized community conversions do
not carry it. This repo preserves it, revision-pinned.
Source: zai-org/GLM-4.7-Flash, revision
7dd20894a642a0aa287e9827cb1a1f7f91386b67 (MIT). All weights are Z.ai's
trained parameters, unmodified except quantization and the layout transforms
below.
Tool: the glm4_moe_lite_mtp drafter split from mlx-vlm's
speculative/drafters convention
(Blaizzy/mlx-vlm#1570):
python -m mlx_vlm.speculative.drafters.glm4_moe_lite_mtp.split \
--model zai-org/GLM-4.7-Flash \
--revision 7dd20894a642a0aa287e9827cb1a1f7f91386b67 \
--output GLM-4.7-Flash-MTP-4bit \
--q-bits 4 --q-group-size 64
Only the 3 (of 48) source shards holding the nextn tensors are read.
Checksum:
| file | sha256 |
|---|---|
model.safetensors | cc2a9750a6328a68b1758502a47f5d286cbc96c26859210fe7c751bbe6d328ba |
model_type: glm4_moe_lite_mtp, block_size: 2
(num_nextn_predict_layers + 1), untied embeddings, affine quantization
(bits: 4, group_size: 64); the source text config is nested under
text_config. 54 tensors, flat post-sanitize layout:
embed_tokens and untied lm_head (GLM's nextn head is not
tied to the target, unlike DeepSeek/Qwen MTP)enorm / hnorm / eh_proj projectionskv_b_proj split into
embed_q / unembed_out)switch_mlp + shared expertnoaux_tc router
correction bias (kept fp32) — casting or quantizing them breaks routingglm4_moe_lite backbone
lands thereMeasured offline acceptance of this head replaying real target hidden states: ~0.806 mean over chat prompts; a 4-bit head forward measured ~6.7% of a target decode step on M3 (methodology and end-to-end results in the vllm-metal links above).
An unquantized variant is at samithaj/GLM-4.7-Flash-MTP-bf16.