Downloads · 30 days
469
0% of all-time downloads
marksverdhei/GLM-4.7-Flash-FP8
GLM-4.7-Flash-FP8 is a text generation model from marksverdhei. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
FP8 quantized version of zai-org/GLM-4.7-Flash.
Downloads · 30 days
469
0% of all-time downloads
All-time downloads
157K
Public
Parameters
31.2B
87.5 GB on disk
Likes
19
Public
Click a slice to open those files.
.safetensors32.2 GB · 100%
How the weights are stored.
F8_E4M330.3B · 97%
From the Hugging Face model README
FP8 quantized version of zai-org/GLM-4.7-Flash.
Also, see Unsloth's new unsloth/GLM-4.7-Flash-FP8-Dynamic
Tested on 2x RTX 3090 (24GB each) with vLLM 0.13.0:
| Setting | Value |
|---|---|
| Tensor Parallel | 2 |
| Context Length | 8192 |
| VRAM per GPU | 14.7 GB |
| Throughput | 19.4 tokens/sec |
Note: RTX 3090 lacks native FP8 support, so vLLM uses the Marlin kernel for weight-only FP8 decompression. GPUs with native FP8 (RTX 40xx, Ada Lovelace+) will achieve higher throughput.
Requires vLLM 0.13.0+ and transformers 5.0+ for glm4_moe_lite architecture support.
from vllm import LLM, SamplingParams
llm = LLM(
model="marksverdhei/GLM-4.7-Flash-fp8",
tensor_parallel_size=2,
max_model_len=8192,
enforce_eager=True, # Optional: disable CUDA graphs to save VRAM
)
outputs = llm.generate(["Hello, world!"], SamplingParams(max_tokens=100))
print(outputs[0].outputs[0].text)
MIT (same as base model)