Downloads · 30 days
2.7K
10% of all-time downloads
RadixArk/Muse-Glimmer-NVFP4
Muse-Glimmer-NVFP4 is a image-text-to-text model from RadixArk. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.
The vendor-recipe quantized checkpoint of Muse Glimmer, packed for serving with SGLang. This is the model vendor's own fpawq mixed-precision recipe (MIXEDPRECISION): NVFP4 (E2M1, group 16) on most projections, MXFP8 (…
Downloads · 30 days
2.7K
10% of all-time downloads
All-time downloads
26.1K
Public
Parameters
18.1B
19.7 GB on disk
Likes
10
Trending 1
Click a slice to open those files.
.safetensors19.7 GB · 100%
How the weights are stored.
U89.8B · 54%
From the Hugging Face model README
The vendor-recipe quantized checkpoint of Muse Glimmer, packed for
serving with SGLang. This is the model vendor's own fp_awq mixed-precision
recipe (MIXED_PRECISION): NVFP4 (E2M1, group 16) on most projections,
MXFP8 (E4M3, group 32) on the higher-sensitivity set (v_proj,
down_proj, lm_head), norms in bf16 — 18.3 GiB versus 52 GiB for the bf16
export. Grab it and serve it directly by repo id; no conversion step needed.
meta-models/Muse-Glimmer-30B (Muse Glimmer bf16 export)fp_awq hand-off packed with the SGLang fork's
convert_fp_awq_to_hf.py (default flags — the recipe as shipped);
quantization map in hf_quant_config.json (quant_algo: MIXED_PRECISION)--language-model-only); the vision
tower ships in the base export, not in this quantized checkpointMeasured on 1× DGX Spark (GB10), SGLang, 1024 in / 1024 out, greedy:
| BS | output tok/s | vs bf16 target | + DFlash draft (sim acc=5) |
|---|---|---|---|
| 1 | 12.1 | 2.7× | 36.4 |
| 4 | 47.7 | 2.7× | 156.2 |
| 8 | 92.2 | 2.6× | 300.7 |
Notes: decode is memory-bound, so the speedup tracks the 52→18 GiB weight reduction. On GB10, quantized prefill is slower than bf16 at batch ≥ 4 (~650 vs ~1500 input tok/s) — decode-heavy workloads win, prefill-heavy workloads should measure. Correctness smoke-verified (greedy arithmetic and generation with correct stop tokens); full accuracy suite pending.
Requires the SGLang fork with Muse Glimmer support (sgl-project/sglang#34262, model support PR #3) until merged upstream.
sglang serve \
--model-path RadixArk/Muse-Glimmer-NVFP4 \
--reasoning-parser muse \
--tool-call-parser muse \
--language-model-only \
--tp-size 1 \
--mem-fraction-static 0.85 \
--host 0.0.0.0 --port 30000
With speculative decoding (pairs with meta-models/Muse-Glimmer-30B-assistant):
sglang serve \
--model-path RadixArk/Muse-Glimmer-NVFP4 \
--reasoning-parser muse \
--tool-call-parser muse \
--language-model-only \
--speculative-algorithm DFLASH \
--speculative-draft-model-path meta-models/Muse-Glimmer-30B-assistant \
--speculative-dflash-block-size 5 \
--tp-size 1 \
--mem-fraction-static 0.85 \
--host 0.0.0.0 --port 30000
Hardware notes:
--mem-fraction-static 0.40
(0.38 with DFlash). Higher fractions let the KV pool consume the shared
CPU/GPU pool and can OOM the machine during load.<|eom|> (200007) as an EOS token — it breaks parallel tool
calling; generation_config.json already carries the correct stops.