Downloads · 30 days
1.1K
1% of all-time downloads
local-inference-lab/GLM-5.2-NVFP4
GLM-5.2-NVFP4 is a text generation model from local-inference-lab. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
GLM-5.2-NVFP4 is an NVFP4-quantized version of zai-org/GLM-5.2, a 744B-parameter Mixture-of-Experts language model with 40B active parameters, 256 experts per MoE layer (8 activated per token), and DeepSeek Sparse Att…
Downloads · 30 days
1.1K
1% of all-time downloads
All-time downloads
122K
Public
Parameters
386B
910 GB on disk
Likes
32
Public
Click a slice to open those files.
.safetensors467 GB · 100%
How the weights are stored.
U8367B · 95%
From the Hugging Face model README
GLM-5.2-NVFP4 is an NVFP4-quantized version of zai-org/GLM-5.2, a 744B-parameter Mixture-of-Experts language model with 40B active parameters, 256 experts per MoE layer (8 activated per token), and DeepSeek Sparse Attention (DSA).
Quantized directly from the full BF16 checkpoint (zai-org/GLM-5.2, not the FP8 release, to NVFP4 (4-bit with blockwise FP8 scales per 16 elements) using NVIDIA Model Optimizer.
Only the non-shared MoE expert MLP projections are quantized to NVFP4. Attention weights are left in BF16, in addition to the dense MLPs (layers 0-3) and the shared experts. Since the MoE expert weights constitute the vast majority of model parameters in an MoE architecture, this still yields significant memory savings.
Calibration uses natural top-k routing rather than forcing all experts to activate, so each expert's quantization scales reflect the token distributions it actually sees during inference. To compensate, calibration was run on a much larger number of samples than typical to ensure broad expert coverage through natural routing alone.
Three calibration passes were run:
Hardware: 8x RTX PRO 6000 Blackwell 96GB (b12x MoE runner recommended)
https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm5.2_v17.md