Downloads · 30 days
72
24% of all-time downloads
SubconsciousDev/glm-5.2-fp8-dflash-v2
glm-5.2-fp8-dflash-v2 is a machine learning model from SubconsciousDev. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This is a DFLASH speculative draft model for GLM-5.2 FP8 serving. The checkpoint uses DFLASH block size 12 and is intended to be loaded as the draft model in SGLang speculative decoding.
Downloads · 30 days
72
24% of all-time downloads
All-time downloads
296
Public
Parameters
2.1B
4.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.2 GB · 100%
From the Hugging Face model README
This is a DFLASH speculative draft model for GLM-5.2 FP8 serving. The checkpoint uses DFLASH block size 12 and is intended to be loaded as the draft model in SGLang speculative decoding.
This model is fine-tuned on top of SubconsciousDev/glm-5.2-fp8-dflash-v1
using SubconsciousDev/Subconscious-Dflash-Training-Dataset-mix-glm52-25k.
Add these arguments to the SGLang launch command:
--speculative-algorithm DFLASH \
--speculative-draft-model-path SubconsciousDev/glm-5.2-fp8-dflash-v2 \
--speculative-num-draft-tokens 12 \
--speculative-draft-kv-cache-dtype bfloat16