Downloads · 30 days
2.1K
41% of all-time downloads
inclusionAI/Ling-3.0-flash-dspark
Ling-3.0-flash-dspark is a text generation model from inclusionAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
A DSpark speculator for Ling3. DSpark extends DFlash with target-model auxiliary features and a confidence head that dynamically chooses the number of draft tokens. The model was trained with SpecForge and is served w…
Downloads · 30 days
2.1K
41% of all-time downloads
All-time downloads
5.1K
Public
Parameters
1.4B
2.7 GB on disk
Likes
21
Trending 1
Click a slice to open those files.
.safetensors2.7 GB · 100%
From the Hugging Face model README
A DSpark speculator for Ling3. DSpark extends DFlash with target-model auxiliary features and a confidence head that dynamically chooses the number of draft tokens. The model was trained with SpecForge and is served with SGLang.
Acceptance length is the mean number of tokens accepted per speculative verification step, including the target bonus token.
| Workload | Acceptance length |
|---|---|
| GSM8K | 6.40 |
| MATH-500 | 6.29 |
| AIME 2025 | 5.56 |
| HumanEval | 6.57 |
| MBPP | 6.34 |
| LiveCodeBench | 5.33 |
| MT-Bench | 3.92 |
| Alpaca | 3.51 |
| Arena-Hard-v2 | 3.72 |
The macro mean across the nine workload means is 5.29.
Launch recipes for this draft on every supported hardware/quantization cell — including the required --linear-replayssm-cache-len sizing — with measured speed and accuracy, are in the SGLang Ling-3.0-flash cookbook.
Use an SGLang version with DSPARK support. Replace the model paths and tensor-parallel size with values appropriate for your deployment:
sglang serve \
--trust-remote-code \
--model-path <LING3_MODEL_PATH> \
--tp-size <TP_SIZE> \
--speculative-algorithm DSPARK \
--speculative-draft-model-path <LING3_DSPARK_MODEL_PATH> \
......
Use a llama.cpp build with DSpark support. Replace the model paths, quantization type, and GPU layer counts with values appropriate for your deployment. First convert and quantize the target model:
python convert_hf_to_gguf.py path/to/Ling-3.0-flash \
--outfile path/to/Ling-3.0-flash-bf16.gguf \
--outtype bf16 --model-name Ling-3.0-flash
llama-quantize \
path/to/Ling-3.0-flash-bf16.gguf \
path/to/Ling-3.0-flash-Q4_K_M.gguf \
Q4_K_M
Then generate the DSpark draft GGUF:
python convert_hf_to_gguf.py \
path/to/Ling-3.0-flash-dspark \
--target-model-dir path/to/Ling-3.0-flash \
--outtype bf16 \
--outfile path/to/Ling-3.0-flash-DSpark.gguf
Finally, launch the server with the DSpark draft as the speculative model:
llama-server \
--model path/to/Ling-3.0-flash-Q4_K_M.gguf \
--spec-draft-model path/to/Ling-3.0-flash-DSpark.gguf \
--spec-type draft-dspark --spec-draft-n-max 8 \
-ngl all -ngld all -fa on