Downloads · 30 days
29
2% of all-time downloads
RESMP-DEV/GLM-4.6-NVFP4
GLM-4.6-NVFP4 is a text generation model from RESMP-DEV. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Quantized version of GLM-4.6 using LLM Compressor and the NVFP4 (E2M1 + E4M3) format.
Downloads · 30 days
29
2% of all-time downloads
All-time downloads
1.7K
Public
Parameters
199B
201 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors201 GB · 100%
How the weights are stored.
U8176B · 88%
From the Hugging Face model README
Quantized version of GLM-4.6 using LLM Compressor and the NVFP4 (E2M1 + E4M3) format.
This time it actually works! We think
This should be the start of a new series of hopefully optimal NVFP4 quantizations as capable cards continue to grow out in the wild.
| Property | Value |
|---|---|
| Base model | GLM-4.6 |
| Quantization | NVFP4 (FP4 microscaling, block = 16, scale = E4M3) |
| Method | Post-Training Quantization with LLM Compressor |
| Toolchain | LLM Compressor |
| Hardware target | NVIDIA Blackwell (Untested on RTX cards) / GB200 Tensor Cores |
| Precision | Weights & activations = FP4 • Scales = FP8 (E4M3) |
| Maintainer | REMSP.DEV |
This model is a drop-in replacement for GLM-4.6 that runs in NVFP4 precision, enabling up to 6× faster GEMM throughput and around 65 % lower memory use compared with BF16. Accuracy remains within ≈ 1 % of the FP8 baseline on standard reasoning and coding benchmarks.