Downloads · 30 days
0
gitcommit90/GLM-5.3-Flash-EXL3-2.05-One-Spark
GLM-5.3-Flash-EXL3-2.05-One-Spark is a machine learning model from gitcommit90. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for vllm. The card lists the license as mit.
A reproducible, production-capable deployment of GLM-5.3-Flash on one NVIDIA DGX Spark, using:
Downloads · 30 days
0
Access
Public
Updated Sep 4, 2026
Repo size
—
Likes
4
Trending 1
Click a slice to open those files.
.md5 KB · 77%
From the Hugging Face model README
A reproducible, production-capable deployment of GLM-5.3-Flash on one NVIDIA DGX Spark, using:
64.1 tok/s structured C1 (K7) · 29.9 tok/s prose / 40.1 tok/s code (K5 default) · 181.9 tok/s C4 active-stream aggregate · 262K context
This is a deployment/runtime discovery page, not a new model or quant. The software, Docker recipe, benchmark harnesses, and raw evidence live in the linked GitHub repository. Target and draft weights download directly from their original publishers.
One DGX Spark, TP1, EXL3 2.05 bpw, DFlash2 K7 (default at collection time; shipped default is now K5), FP8 KV, thinking disabled:
| Benchmark | Result |
|---|---|
| Structured C1, five-run median, temperature 0 | 64.053 tok/s |
| Structured C1, five-run median, temperature 1.0/top-p 0.95 | 62.637 tok/s |
| Open-ended prose, five-run median | 25.059 tok/s |
| C4 median active stream | 40.958 tok/s/stream |
| C4 summed active-stream convention | 181.944 tok/s |
| C4 strict submission-to-completion wall | 91.475 tok/s |
C4 disclosure: 181.944 tok/s is the sum of each stream's active decode rate, matching MiaAI's reporting convention. Because the current scheduler stages admission, strict full-batch wall throughput is 91.475 tok/s. Both numbers are published intentionally.
Default num_speculative_tokens moved from 7 to 5 on 2026-09-03. Spec decode is lossless; K only changes speed. Temperature 0, thinking off, 400 tokens, C1:
| K | Structured | Prose | Code |
|---|---|---|---|
| 4 | 48.6 tok/s | 28.8 | 37.7 |
| 5 (default) | 53.5 | 29.9 | 40.1 |
| 6 | 59.2 | 29.0 | 37.5 |
| 7 | 63.9 | 25.8 | 38.3 |
| 8 | 66.9 | 23.9 | 38.0 |
Use K=8 for list/JSON-heavy workloads (ONE_SPARK_K=8). Raw data in the GitHub repo under benchmarks/raw/k-sweep-20260903.
| Prompt | Cold prefill | Cold TTFT | Warm TTFT |
|---|---|---|---|
| 8K | 786.4 tok/s | 10.175 s | 1.775 s |
| 16K | 822.1 tok/s | 19.464 s | 2.786 s |
| 100K | 845.7 tok/s | 118.250 s | 8.842 s |
Every cold request had zero prefix-cache hits; every warm checksum response was correct.
This deployment uses Turboderp's GLM-5.3-Flash EXL3 2.05-bpw quant, revision 51058cd551c7e570d87bd32a4adee720edce2349. The exact checkpoint is 85.23 GB (79.38 GiB).
An independent full-vocabulary measurement of this exact revision against BF16 teacher logits reported:
| Quant | Size | Top-1 agreement | Mean KLD | Scored positions |
|---|---|---|---|---|
| Turboderp EXL3 2.05 | 85.23 GB | 88.92% | 0.121638 | 51,175 |
The measurement used 25 windows, the full 154,880-token vocabulary, teacher forcing, FP64 accumulation, and two cold runs with identical results. The checkpoint was quantized from the official FP8 release; the measurement reference is BF16. Machine-readable summary: quantization-analysis.json.
https://github.com/gitcommit90/glm-5.3-one-sparkghcr.io/gitcommit90/glm-5.3-one-spark:general23The runtime is derived from Mia's AI Lab's MIT-licensed two-Spark recipe and substantially adapted for TP1, full-model 2.05-bpw mul1 EXL3, ARM64/SM121, and one-Spark memory limits.
Credits: Z.ai / GLM-5 Team, Turboderp and ExLlamaV3, Inco AI, Mia's AI Lab, and the vLLM contributors.
DFlash2 is CC BY-NC-ND 4.0 for research/evaluation. It is not bundled here. Commercial users must obtain appropriate licensing from Inco AI.