Downloads · 30 days
37
19% of all-time downloads
0xSero/GLM-4.7-Flash
GLM-4.7-Flash is a text generation model from 0xSero. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
[!TIP] Support this work → · X · GitHub · REAP paper · Cerebras REAP
Downloads · 30 days
37
19% of all-time downloads
All-time downloads
190
Public
Parameters
29.9B
120 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors59.9 GB · 100%
How the weights are stored.
BF1629.9B · 100%
From the Hugging Face model README
[!TIP] Support this work → · X · GitHub · REAP paper · Cerebras REAP
REAP-pruned zai-org/GLM-4.7-Flash.
| Base model | zai-org/GLM-4.7-Flash |
| Format | BF16 |
| Total params | 30B |
| Active / token | — |
| Experts / layer | 64 |
| Layers | 47 |
| Hidden size | 2048 |
| Context | 202,752 |
| On-disk size | 120 GB |
| Variant | Format | Link |
|---|---|---|
GLM-4.7-Flash (this) | BF16 | link |
GLM-4.7-Flash-DPO | DPO | link |
GLM-4.7-Flash-SFT | SFT | link |
GLM-4.7-Flash-Tools | Tools | link |
License inherited from the base model.
@misc{lasby2025reap,
title = {REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression},
author = {Mike Lasby and Ivan Lazarevich and Nish Sinnadurai and Sean Lie and Yani Ioannou and Vithursan Thangarasa},
year = {2025}, eprint = {2510.13999}, archivePrefix = {arXiv}
}
Made possible by NVIDIA · TNG Technology · Lambda · Prime Intellect · Hot Aisle.