Downloads · 30 days
113
1% of all-time downloads
protoLabsAI/Qwythos-9B-v2-NVFP4
Qwythos-9B-v2-NVFP4 is a text generation model from protoLabsAI. Use it when you need the model to write or continue text. It is set up for vllm. The card lists the license as apache-2.0.
NVFP4 (W4A4) build of empero-ai/Qwythos-9B-v2 for vLLM on NVIDIA Blackwell — the native 4-bit serving path nobody ships for this checkpoint. Empero already publishes the MTP-GGUF for llama.cpp; this is the piece that…
Downloads · 30 days
113
1% of all-time downloads
All-time downloads
10K
Public
Parameters
7.1B
11.7 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors11.7 GB · 100%
How the weights are stored.
BF164.1B · 58%
From the Hugging Face model README
NVFP4 (W4A4) build of empero-ai/Qwythos-9B-v2 for vLLM on NVIDIA Blackwell — the native 4-bit serving path nobody ships for this checkpoint. Empero already publishes the MTP-GGUF for llama.cpp; this is the piece that was missing: a compressed-tensors NVFP4 artifact that runs the model's GEMMs directly on the Blackwell FP4 tensor cores under vLLM.
Quantized by protoLabs. Base model, its capabilities, and its uncensored research posture are Empero AI's — see their card.
linear_attn) layers, the vision tower, lm_head, and the MTP head are kept BF16 — DeltaNet corrupts under 4-bit, and the vision path stays lossless.model-mtp.safetensors) — see the status note below.Single RTX PRO 6000 Blackwell (sm120), vLLM 0.24.0, NVFP4 --linear-backend marlin, BF16 KV:
| Regime | tok/s | TPOT |
|---|---|---|
| single-stream (real text) | 138 | — |
| chat 1k/1k · C1 | 152 | 6.5 ms |
| chat 1k/1k · C8 | 1051 | 7.1 ms |
| chat 1k/1k · C32 | 2523 | 11.4 ms |
(Client-side vllm bench serve, random dataset, fixed seed. --ignore-eos. Single trial.)
vllm serve protoLabsAI/Qwythos-9B-v2-NVFP4 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--linear-backend marlin \
--max-model-len 65536 \
--trust-remote-code
sm120 requires the FlashInfer/CUDA-13 env (or you hit the "no CUDA arch for major 12" crash):
VLLM_USE_FLASHINFER_SAMPLER=0, CUDA_HOME=<cu13>, FLASHINFER_CUDA_ARCH_LIST=12.0f,
NVCC_APPEND_FLAGS=-DCCCL_DISABLE_CTK_COMPATIBILITY_CHECK. --linear-backend marlin is the
proven NVFP4 kernel on sm120 (fastest at the single-stream / low-batch regime that matters here).
llm-compressor NVFP4, calibrated (512 samples, seq 2048). Ignore list: lm_head, re:.*visual.*,
re:.*linear_attn.*, re:.*mtp.*. VL keys canonicalized post-quant. Reproducible from the BF16 source.
Every suite run against both the NVFP4 build and the BF16 source, same prompts/harness, on this rig. Deterministic suites are judge-free (solver-verified reasoning, execution-graded code); claw uses an independent LLM grader. Single trial (coding ×3).
| Suite | NVFP4 | BF16 base | Δ |
|---|---|---|---|
| Function-calling | 89% (48/54) | 94% (51/54) | −5 pp |
| Reasoning-v2 (solver-verified) | 0.759 | 0.789 | −0.03 |
| Coding (exec-graded, ×3) | 0.518 | 0.553 | −0.03 |
| Claw (agentic, 10 tasks) | 0.651 | 0.614 | +0.04 |
| Coherence @ depth 4K–60K | clean | — | needle ✓, no repetition/degradation |
Near-parity. The only real cost is a few function-calling cases (single-trial, 54-task suite); reasoning and coding are within noise, agentic is flat-to-better, and there is no quant-rot at long context (verified to 60K). Expected NVFP4 behaviour for a 9B.
Need a different size or format (GGUF, other quant, longer-context config)? Open a Community discussion — we usually ship within 48h.
apache-2.0, inherited from the base. Uncensored — this is Empero's research/red-team model; the quant changes numerics, not alignment. Use responsibly and within the base model's terms.